Zhenhua Liu 0003

dblp:02/1825-3 · DBLP profile ↗
← Back
10ranked-venue papers
4as first author
9since 2021 · last 2024
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 8 · 3 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 2 first-author · 6 since 2021
YearPublicationVenuePosition
2024 Kangaroo: Lossless Self-Speculative Decoding for Accelerating LLMs via Double Early Exiting
abstract
Speculative decoding has demonstrated its effectiveness in accelerating the inference of large language models (LLMs) while maintaining an identical sampling distribution. However, the conventional approach of training separate draft model to achieve a satisfactory token acceptance rate can be costly and impractical. In this paper, we propose a novel self-speculative decoding framework \emph{Kangaroo} with \emph{double} early exiting strategy, which leverages the shallow sub-network and the \texttt{LM Head} of the well-trained target LLM to construct a self-drafting model. Then, the self-verification stage only requires computing the remaining layers over the \emph{early-exited} hidden states in parallel. To bridge the representation gap between the sub-network and the full model, we train a lightweight and efficient adapter module on top of the sub-network. One significant challenge that comes with the proposed method is that the inference latency of the self-draft model may no longer be negligible compared to the big model. To boost the token acceptance rate while minimizing the latency of the self-drafting model, we introduce an additional \emph{early exiting} mechanism for both single-sequence and the tree decoding scenarios. Specifically, we dynamically halt the small model's subsequent prediction during the drafting phase once the confidence level for the current step falls below a certain threshold. This approach reduces unnecessary computations and improves overall efficiency. Extensive experiments on multiple benchmarks demonstrate our effectiveness, where Kangaroo achieves walltime speedups up to 2.04$\times$, outperforming Medusa-1 with 88.7\% fewer additional parameters. The code for Kangaroo is available at https://github.com/Equationliu/Kangaroo.
Fangcheng Liu, Yehui Tang 0001, Zhenhua Liu 0003, Yunsheng Ni, Duyu Tang, Kai Han 0002, Yunhe Wang 0001
NeurIPS3
2023 Diffusion-Based 3D Human Pose Estimation with Multi-Hypothesis Aggregation
abstract
In this paper, a novel Diffusion-based 3D Pose estimation (D3DP) method with Joint-wise reProjection-based Multi-hypothesis Aggregation (JPMA) is proposed for probabilistic 3D human pose estimation. On the one hand, D3DP generates multiple possible 3D pose hypotheses for a single 2D observation. It gradually diffuses the ground truth 3D poses to a random distribution, and learns a denoiser conditioned on 2D keypoints to recover the uncontaminated 3D poses. The proposed D3DP is compatible with existing 3D pose estimators and supports users to balance efficiency and accuracy during inference through two customizable parameters. On the other hand, JPMA is proposed to assemble multiple hypotheses generated by D3DP into a single 3D pose for practical use. It reprojects 3D pose hypotheses to the 2D camera plane, selects the best hypothesis joint-by-joint based on the reprojection errors, and combines the selected joints into the final pose. The proposed JPMA conducts aggregation at the joint level and makes use of the 2D prior information, both of which have been overlooked by previous approaches. Extensive experiments on Human3.6M and MPI-INF-3DHP datasets show that our method outperforms the state-of-the-art deterministic and probabilistic approaches by 1.5% and 8.9%, respectively. Code is available at https://github.com/paTRICK-swk/D3DP.
Wenkang Shan, Zhenhua Liu 0003, Xinfeng Zhang 0001, Zhao Wang 0004, Kai Han 0002, Shanshe Wang, Siwei Ma 0001, Wen Gao 0001
ICCV2
2023 A Survey on Vision Transformer
abstract
Transformer, first applied to the field of natural language processing, is a type of deep neural network mainly based on the self-attention mechanism. Thanks to its strong representation capabilities, researchers are looking at ways to apply transformer to computer vision tasks. In a variety of visual benchmarks, transformer-based models perform similar to or better than other types of networks such as convolutional and recurrent neural networks. Given its high performance and less need for vision-specific inductive bias, transformer is receiving more and more attention from the computer vision community. In this paper, we review these vision transformer models by categorizing them in different tasks and analyzing their advantages and disadvantages. The main categories we explore include the backbone network, high/mid-level vision, low-level vision, and video processing. We also include efficient transformer methods for pushing transformer into real device-based applications. Furthermore, we also take a brief look at the self-attention mechanism in computer vision, as it is the base component in transformer. Toward the end of this paper, we discuss the challenges and provide several further research directions for vision transformers.
Kai Han 0002, Yunhe Wang 0001, Hanting Chen, Xinghao Chen 0001, Jianyuan Guo, Zhenhua Liu 0003, Yehui Tang 0001, An Xiao, Chunjing Xu, Yixing Xu, Zhaohui Yang 0003, Dacheng Tao
IEEE Trans. Pattern Anal. Mach. Intell.6
2023 Differential Weight Quantization for Multi-Model Compression
abstract
Low bit-width quantization can effectively reduce the storage and computational costs of deep neural networks. Existing quantization methods are commonly designed for single model compression. For multi-model compression scenarios, multiple models for the same task or similar tasks need to be compressed simultaneously in multimedia tasks, such as compressing image super-resolution models for different scales and transferring of different models in multimedia. However, single model quantization methods do not consider the correlations among the weights of different models, which limits the further compression for the above multi-model compression scenarios. To sufficiently excavate the potential of compression on multi-model, we propose a novel quantization scheme for multi-model compression, namely differential weight quantization (DWQ), which focuses on the weights increment between the target model and the reference model. Specifically, DWQ is achieved by increment computation, increment quantization and fine-tuning, which utilizes the reference model to guide the subsequent quantization on the target model. Due to the correlations between the weights of different models, the distribution of weights increment is more centralized compared with original weights, which can achieve a higher compression ratio by lower bit representation on weights increment. Moreover, the progressive training method is proposed to accelerate the convergence and reduce quantization loss on the DWQ framework. Extensive experiments validate the effectiveness of DWQ based on weight-sharing and parameterized clipping activation (PACT) quantization technologies on multiple tasks. The proposed framework can achieve 2× compression improvement and reduce 30% computational complexity with comparable performance in the popular multimedia tasks.
Wenhong Duan, Zhenhua Liu 0003, Chuanmin Jia, Shanshe Wang, Siwei Ma 0001, Wen Gao 0001
IEEE Trans. Multim.2
2022 Instance-Aware Dynamic Neural Network Quantization
abstract
Quantization is an effective way to reduce the memory and computational costs of deep neural networks in which the full-precision weights and activations are represented using low-bit values. The bit-width for each layer in most of existing quantization methods is static, i.e., the same for all samples in the given dataset. However, natural images are of huge diversity with abundant content and using such a universal quantization configuration for all samples is not an optimal strategy. In this paper, we present to conduct the low-bit quantization for each image individually, and develop a dynamic quantization scheme for exploring their optimal bit-widths. To this end, a lightweight bit-controller is established and trained jointly with the given neural network to be quantized. During inference, the quantization configuration for an arbitrary image will be determined by the bit-widths generated by the controller, e.g., an image with simple texture will be allocated with lower bits and computational complexity and vice versa. Experimental results conducted on benchmarks demonstrate the effectiveness of the proposed dynamic quantization method for achieving state-of-art performance in terms of accuracy and computational complexity. The code will be available at https://github.com/huawei-noah/Efficient-Computing and https://gitee.com/mindspore/models/tree/master/research/cv/DynamicQuant.
Zhenhua Liu 0003, Yunhe Wang 0001, Kai Han 0002, Siwei Ma 0001, Wen Gao 0001
CVPR1
2022 P-STMO: Pre-trained Spatial Temporal Many-to-One Model for 3D Human Pose Estimation
Wenkang Shan, Zhenhua Liu 0003, Xinfeng Zhang 0001, Shanshe Wang, Siwei Ma 0001, Wen Gao 0001
ECCV (5)2
2021 Pre-Trained Image Processing Transformer
abstract
As the computing power of modern hardware is increasing strongly, pre-trained deep learning models (e.g., BERT, GPT-3) learned on large-scale datasets have shown their effectiveness over conventional methods. The big progress is mainly contributed to the representation ability of transformer and its variant architectures. In this paper, we study the low-level computer vision task (e.g., denoising, super-resolution and deraining) and develop a new pre-trained model, namely, image processing transformer (IPT). To maximally excavate the capability of transformer, we present to utilize the well-known ImageNet benchmark for generating a large amount of corrupted image pairs. The IPT model is trained on these images with multi-heads and multi-tails. In addition, the contrastive learning is introduced for well adapting to different image processing tasks. The pre-trained model can therefore efficiently employed on desired task after fine-tuning. With only one pre-trained model, IPT outperforms the current state-of-the-art methods on various low-level benchmarks. Code is available at https://github.com/huawei-noah/Pretrained-IPT and https://gitee.com/mindspore/mindspore/tree/master/model_zoo/research/cv/IPT
Hanting Chen, Yunhe Wang 0001, Tianyu Guo 0001, Chang Xu 0002, Yiping Deng, Zhenhua Liu 0003, Siwei Ma 0001, Chunjing Xu, Chao Xu 0006, Wen Gao 0001
CVPR6
2021 Evolutionary Quantization of Neural Networks with Mixed-Precision
abstract
Quantization is an effective way for reducing the memory and computation costs of deep neural networks. Most of existing methods exploit the fixed-precision quantization approach, e.g., weights and activations (i.e., output features) are represented as 8-bit values. Although mixed-precision quantization provides us a greater possibility to efficiently allocate computation resources and maintain the network performance, it is difficult to accurately solve the optimal bit-width of each layer. In this paper, we develop a novel evolutionary based method to automatically determine the bit-widths of weights and activations in each convolutional layer, namely, Evolutionary Mixed-Precision Quantization (EMQ). Specifically, the quantization intervals of weights and activations of all layers in the given network will be simultaneously encoded as an individual. The fitness of each individual is calculated as the performance of the corresponding quantized network. The optimal quantization result will be updated and elected during the evolutionary search. Extensive experiments conducted on benchmark datasets and models demonstrate the effectiveness of the proposed method over the state-of-the-art network quantization algorithms.
Zhenhua Liu 0003, Xinfeng Zhang 0001, Shanshe Wang, Siwei Ma 0001, Wen Gao 0001
ICASSP1
2021 Post-Training Quantization for Vision Transformer
abstract
Recently, transformer has achieved remarkable performance on a variety of computer vision applications. Compared with mainstream convolutional neural networks, vision transformers are often of sophisticated architectures for extracting powerful feature representations, which are more difficult to be developed on mobile devices. In this paper, we present an effective post-training quantization algorithm for reducing the memory storage and computational costs of vision transformers. Basically, the quantization task can be regarded as finding the optimal low-bit quantization intervals for weights and inputs, respectively. To preserve the functionality of the attention mechanism, we introduce a ranking loss into the conventional quantization objective that aims to keep the relative order of the self-attention results after quantization. Moreover, we thoroughly analyze the relationship between quantization loss of different layers and the feature diversity, and explore a mixed-precision quantization scheme by exploiting the nuclear norm of each attention map and output feature. The effectiveness of the proposed method is verified on several benchmark models and datasets, which outperforms the state-of-the-art post-training quantization algorithms. For instance, we can obtain an 81.29% top-1 accuracy using DeiT-B model on ImageNet dataset with about 8-bit quantization. Code will be available at https://gitee.com/mindspore/models/tree/master/research/cv/VT-PTQ.
Zhenhua Liu 0003, Yunhe Wang 0001, Kai Han 0002, Wei Zhang 0196, Siwei Ma 0001, Wen Gao 0001
NeurIPS1
2018 Frequency-Domain Dynamic Pruning for Convolutional Neural Networks
abstract
Deep convolutional neural networks have demonstrated their powerfulness in a variety of applications. However, the storage and computational requirements have largely restricted their further extensions on mobile devices. Recently, pruning of unimportant parameters has been used for both network compression and acceleration. Considering that there are spatial redundancy within most filters in a CNN, we propose a frequency-domain dynamic pruning scheme to exploit the spatial correlations. The frequency-domain coefficients are pruned dynamically in each iteration and different frequency bands are pruned discriminatively, given their different importance on accuracy. Experimental results demonstrate that the proposed scheme can outperform previous spatial-domain counterparts by a large margin. Specifically, it can achieve a compression ratio of 8.4x and a theoretical inference speed-up of 9.2x for ResNet-110, while the accuracy is even better than the reference model on CIFAR-110.
Zhenhua Liu 0003, Jizheng Xu, Xiulian Peng, Ruiqin Xiong
NeurIPS1