Jun Chen 0023

dblp:85/5901-23 · DBLP profile ↗
← Back
32ranked-venue papers
5as first author
32since 2021 · last 2026
0000-0001-6568-8801ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 30 · 5 first-author · 30 since 2021Graphics, computer vision, multimedia, augmented reality and games · 9 · 9 since 2021
YearPublicationVenuePosition
2026 Automatic data-free pruning via channel similarity reconstruction
Siqi Li 0009, Jun Chen 0023, Jingyang Xiang, Chengrui Zhu, Jiandang Yang, Xiaobin Wei, Yunliang Jiang, Yong Liu 0007
Neurocomputing2
2026 SampleLLM: Prune LLMs via learnable structure sampling
Jiateng Wei, Siqi Li 0009, Jingyang Xiang, Longqi Wang, Chuang Guo, Jun Chen 0023, Baochang Zhu, Yong Liu 0007
Neurocomputing7
2026 WLR: Well-conditioned linear reconstruction for retraining-free pruning of LLMs
Siqi Li 0009, Jingyang Xiang, Jiateng Wei, Chengrui Zhu, Jiandang Yang, Jun Chen 0023, Jian Yang 0003, Xiaobin Wei, Yunliang Jiang, Yong Liu 0007
Neural Networks6
2026 OOPS: Outlier-aware and quadratic programming based structured pruning for large language models
Jiateng Wei, Siqi Li 0009, Jingyang Xiang, Jiandang Yang, Jun Chen 0023, Xiaobin Wei, Yunliang Jiang, Yong Liu 0007
Neural Networks5
2025 Manifold Constraint Reduces Exposure Bias in Accelerated Diffusion Sampling
abstract
Diffusion models have demonstrated significant potential for generating high-quality images, audio, and videos. However, their iterative inference process entails substantial computational costs, limiting practical applications. Recently, researchers have introduced accelerated sampling methods that enable diffusion models to generate samples with far fewer timesteps than those used during training. Nonetheless, as the number of sampling steps decreases, the prediction errors significantly degrade the quality of generated outputs. Additionally, the exposure bias in diffusion models further amplifies these errors. To address these challenges, we leverage a manifold hypothesis to explore the exposure bias problem in depth. Based on this geometric perspective, we propose a manifold constraint that effectively reduces exposure bias during accelerated sampling of diffusion models. Notably, our method involves no additional training and requires only minimal hyperparameter tuning. Extensive experiments demonstrate the effectiveness of our approach, achieving a FID score of 15.60 with 10-step SDXL on MS-COCO, surpassing the baseline by a reduction of 2.57 in FID.
Yuzhe Yao, Jun Chen 0023, Zeyi Huang, Haonan Lin, Mengmeng Wang 0005, Guang Dai, Jingdong Wang 0001
ICLR2
2025 Hyperbolic Binary Neural Network
abstract
Binary neural network (BNN) converts full-precision weights and activations into their extreme 1-bit counterparts, making it particularly suitable for deployment on lightweight mobile devices. While BNNs are typically formulated as a constrained optimization problem and optimized in the binarized space, general neural networks are formulated as an unconstrained optimization problem and optimized in the continuous space. This article introduces the hyperbolic BNN (HBNN) by leveraging the framework of hyperbolic geometry to optimize the constrained problem. Specifically, we transform the constrained problem in hyperbolic space into an unconstrained one in Euclidean space using the Riemannian exponential map. On the other hand, we also propose the exponential parametrization cluster (EPC) method, which, compared with the Riemannian exponential map, shrinks the segment domain based on a diffeomorphism. This approach increases the probability of weight flips, thereby maximizing the information gain in BNNs. Experimental results on CIFAR10, CIFAR100, and ImageNet classification datasets with VGGsmall, ResNet18, and ResNet34 models illustrate the superior performance of our HBNN over state-of-the-art methods.
Jun Chen 0023, Jingyang Xiang, Tianxin Huang, Xiangrui Zhao, Yong Liu 0007
IEEE Trans. Neural Networks Learn. Syst.1
2025 MCMC: Multi-Constrained Model Compression via One-Stage Envelope Reinforcement Learning
abstract
Model compression methods are being developed to bridge the gap between the massive scale of neural networks and the limited hardware resources on edge devices. Since most real-world applications deployed on resource-limited hardware platforms typically have multiple hardware constraints simultaneously, most existing model compression approaches that only consider optimizing one single hardware objective are ineffective. In this article, we propose an automated pruning method called multi-constrained model compression (MCMC) that allows for the optimization of multiple hardware targets, such as latency, floating point operations (FLOPs), and memory usage, while minimizing the impact on accuracy. Specifically, we propose an improved multi-objective reinforcement learning (MORL) algorithm, the one-stage envelope deep deterministic policy gradient (DDPG) algorithm, to determine the pruning strategy for neural networks. Our improved one-stage envelope DDPG algorithm reduces exploration time and offers greater flexibility in adjusting target priorities, enhancing its suitability for pruning tasks. For instance, on the visual geometry group (VGG)-16 network, our method achieved an 80% reduction in FLOPs, a reduction in memory usage, and a acceleration, with an accuracy improvement of 0.09% compared with the baseline. For larger datasets, such as ImageNet, we reduced FLOPs by 50% for MobileNet-V1, resulting in a faster speed and memory compression, while maintaining the same accuracy. When applied to edge devices, such as JETSON XAVIER NX, our method resulted in a 71% reduction in FLOPs for MobileNet-V1, leading to a faster speed, memory compression, and an accuracy improvement.
Siqi Li 0009, Jun Chen 0023, Shanqi Liu, Chengrui Zhu, Guanzhong Tian, Yong Liu 0007
IEEE Trans. Neural Networks Learn. Syst.2
2025 Adding Before Pruning: Sparse Filter Fusion for Deep Convolutional Neural Networks via Auxiliary Attention
abstract
Filter pruning is a significant feature selection technique to shrink the existing feature fusion schemes (especially on convolution calculation and model size), which helps to develop more efficient feature fusion models while maintaining state-of-the-art performance. In addition, it reduces the storage and computation requirements of deep neural networks (DNNs) and accelerates the inference process dramatically. Existing methods mainly rely on manual constraints such as normalization to select the filters. A typical pipeline comprises two stages: first pruning the original neural network and then fine-tuning the pruned model. However, choosing a manual criterion can be somehow tricky and stochastic. Moreover, directly regularizing and modifying filters in the pipeline suffer from being sensitive to the choice of hyperparameters, thus making the pruning procedure less robust. To address these challenges, we propose to handle the filter pruning issue through one stage: using an attention-based architecture that adaptively fuses the filter selection with filter learning in a unified network. Specifically, we present a pruning method named adding before pruning (ABP) to make the model focus on the filters of higher significance by training instead of man-made criteria such as norm, rank, etc. First, we add an auxiliary attention layer into the original model and set the significance scores in this layer to be binary. Furthermore, to propagate the gradients in the auxiliary attention layer, we design a specific gradient estimator and prove its effectiveness for convergence in the graph flow through mathematical derivation. In the end, to relieve the dependence on the complicated prior knowledge for designing the thresholding criterion, we simultaneously prune and train the filters to automatically eliminate network redundancy with recoverability. Extensive experimental results on the two typical image classification benchmarks, CIFAR-10 and ILSVRC-2012, illustrate that the proposed approach performs favorably against previous state-of-the-art filter pruning algorithms.
Guanzhong Tian, Yiran Sun, Yuang Liu, Xianfang Zeng, Mengmeng Wang 0005, Yong Liu 0007, Jiangning Zhang, Jun Chen 0023
IEEE Trans. Neural Networks Learn. Syst.8
2024 A Multimodal, Multi-Task Adapting Framework for Video Action Recognition
abstract
Recently, the rise of large-scale vision-language pretrained models like CLIP, coupled with the technology of Parameter-Efficient FineTuning (PEFT), has captured substantial attraction in video action recognition. Nevertheless, prevailing approaches tend to prioritize strong supervised performance at the expense of compromising the models' generalization capabilities during transfer. In this paper, we introduce a novel Multimodal, Multi-task CLIP adapting framework named M2-CLIP to address these challenges, preserving both high supervised performance and robust transferability. Firstly, to enhance the individual modality architectures, we introduce multimodal adapters to both the visual and text branches. Specifically, we design a novel visual TED-Adapter, that performs global Temporal Enhancement and local temporal Difference modeling to improve the temporal representation capabilities of the visual encoder. Moreover, we adopt text encoder adapters to strengthen the learning of semantic label information. Secondly, we design a multi-task decoder with a rich set of supervisory signals, including the original contrastive learning head, a cross-modal classification head, a cross-modal masked language modeling head, and a visual classification head. This multi-task decoder adeptly satisfies the need for strong supervised performance within a multimodal framework. Experimental results validate the efficacy of our approach, demonstrating exceptional performance in supervised learning while maintaining strong generalization in zero-shot scenarios.
Mengmeng Wang 0005, Jiazheng Xing, Boyuan Jiang, Jun Chen 0023, Jianbiao Mei, Xingxing Zuo 0001, Guang Dai, Jingdong Wang 0001, Yong Liu 0007
AAAI4
2024 Timestep-Aware Correction for Quantized Diffusion Models
Yuzhe Yao, Feng Tian 0002, Jun Chen 0023, Haonan Lin, Guang Dai, Yong Liu 0007, Jingdong Wang 0001
ECCV (66)3
2024 Structured Optimal Brain Pruning for Large Language Models
abstract
The massive parameters and computational demands hinder the widespread application of Large Language Models (LLMs).Network pruning provides a practical solution to this problem.However, existing pruning works for LLMs mainly focus on unstructured pruning or necessitate post-pruning fine-tuning.The former relies on special hardware to accelerate computation, while the latter may need substantial computational resources.In this paper, we introduce a retraining-free structured pruning method called SoBP (Structured Optimal Brain Pruning).It leverages global first-order information to select pruning structures, then refines them with a local greedy approach, and finally adopts module-wise reconstruction to mitigate information loss.We assess the effectiveness of SoBP across 14 models from 3 LLM families on 8 distinct datasets.Experimental results demonstrate that SoBP outperforms current state-of-the-art methods.
Jiateng Wei, Siqi Li 0009, Jingyang Xiang, Jun Chen 0023, Yong Liu 0007
EMNLP6
2024 Decentralized Riemannian Conjugate Gradient Method on the Stiefel Manifold
abstract
The conjugate gradient method is a crucial first-order optimization method that generally converges faster than the steepest descent method, and its computational cost is much lower than that of second-order methods. However, while various types of conjugate gradient methods have been studied in Euclidean spaces and on Riemannian manifolds, there is little study for those in distributed scenarios. This paper proposes a decentralized Riemannian conjugate gradient descent (DRCGD) method that aims at minimizing a global function over the Stiefel manifold. The optimization problem is distributed among a network of agents, where each agent is associated with a local function, and the communication between agents occurs over an undirected connected graph. Since the Stiefel manifold is a non-convex set, a global function is represented as a finite sum of possibly non-convex (but smooth) local functions. The proposed method is free from expensive Riemannian geometric operations such as retractions, exponential maps, and vector transports, thereby reducing the computational complexity required by each agent. To the best of our knowledge, DRCGD is the first decentralized Riemannian conjugate gradient algorithm to achieve global convergence over the Stiefel manifold.
Jun Chen 0023, Haishan Ye, Mengmeng Wang 0005, Tianxin Huang, Guang Dai, Ivor W. Tsang, Yong Liu 0007
ICLR1
2024 Towards efficient filter pruning via adaptive automatic structure search
Xiaozhou Xu, Jun Chen 0023, Zhishan Li, Lei Xie 0007
Eng. Appl. Artif. Intell.2
2024 Learning Discretized Neural Networks under Ricci Flow
abstract
In this paper, we study Discretized Neural Networks (DNNs) composed of low-precision weights and activations, which suffer from either infinite or zero gradients due to the non-differentiable discrete function during training. Most training-based DNNs in such scenarios employ the standard Straight-Through Estimator (STE) to approximate the gradient w.r.t. discrete values. However, the use of STE introduces the problem of gradient mismatch, arising from perturbations in the approximated gradient. To address this problem, this paper reveals that this mismatch can be interpreted as a metric perturbation in a Riemannian manifold, viewed through the lens of duality theory. Building on information geometry, we construct the Linearly Nearly Euclidean (LNE) manifold for DNNs, providing a background for addressing perturbations. By introducing a partial differential equation on metrics, i.e., the Ricci flow, we establish the dynamical stability and convergence of the LNE metric with the $L^2$-norm perturbation. In contrast to previous perturbation theories with convergence rates in fractional powers, the metric perturbation under the Ricci flow exhibits exponential decay in the LNE manifold. Experimental results across various datasets demonstrate that our method achieves superior and more stable performance for DNNs compared to other representative training-based methods.
Jun Chen 0023, Hanwen Chen, Mengmeng Wang 0005, Guang Dai, Ivor W. Tsang, Yong Liu 0007
J. Mach. Learn. Res.1
2024 Semi-supervised noise-resilient anomaly detection with feature autoencoder
Lina Liu 0010, Yuanlong Zhang, Chao Xu 0023, Jun Chen 0023
Knowl. Based Syst.7
2024 Learnable Chamfer Distance for point cloud reconstruction
Tianxin Huang, Qingyao Liu, Xiangrui Zhao, Jun Chen 0023, Yong Liu 0007
Pattern Recognit. Lett.4
2024 DCCD: Reducing Neural Network Redundancy via Distillation
abstract
Deep neural models have achieved remarkable performance on various supervised and unsupervised learning tasks, but it is a challenge to deploy these large-size networks on resource-limited devices. As a representative type of model compression and acceleration methods, knowledge distillation (KD) solves this problem by transferring knowledge from heavy teachers to lightweight students. However, most distillation methods focus on imitating the responses of teacher networks but ignore the information redundancy of student networks. In this article, we propose a novel distillation framework difference-based channel contrastive distillation (DCCD), which introduces channel contrastive knowledge and dynamic difference knowledge into student networks for redundancy reduction. At the feature level, we construct an efficient contrastive objective that broadens student networks' feature expression space and preserves richer information in the feature extraction stage. At the final output level, more detailed knowledge is extracted from teacher networks by making a difference between multiview augmented responses of the same instance. We enhance student networks to be more sensitive to minor dynamic changes. With the improvement of two aspects of DCCD, the student network gains contrastive and difference knowledge and reduces its overfitting and redundancy. Finally, we achieve surprising results that the student approaches and even outperforms the teacher in test accuracy on CIFAR-100. We reduce the top-1 error to 28.16% on ImageNet classification and 24.15% for cross-model transfer with ResNet-18. Empirical experiments and ablation studies on popular datasets show that our proposed method can achieve state-of-the-art accuracy compared with other distillation methods.
Yuang Liu, Jun Chen 0023, Yong Liu 0007
IEEE Trans. Neural Networks Learn. Syst.2
2024 Learning Multi-Agent Cooperation via Considering Actions of Teammates
abstract
Recently value-based centralized training with decentralized execution (CTDE) multi-agent reinforcement learning (MARL) methods have achieved excellent performance in cooperative tasks. However, the most representative method among these methods, Q-network MIXing (QMIX), restricts the joint action Q values to be a monotonic mixing of each agent's utilities. Furthermore, current methods cannot generalize to unseen environments or different agent configurations, which is known as ad hoc team play situation. In this work, we propose a novel Q values decomposition that considers both the return of an agent acting on its own and cooperating with other observable agents to address the nonmonotonic problem. Based on the decomposition, we propose a greedy action searching method that can improve exploration and is not affected by changes in observable agents or changes in the order of agents' actions. In this way, our method can adapt to ad hoc team play situation. Furthermore, we utilize an auxiliary loss related to environmental cognition consistency and a modified prioritized experience replay (PER) buffer to assist training. Our extensive experimental results show that our method achieves significant performance improvements in both challenging monotonic and nonmonotonic domains, and can handle the ad hoc team play situation perfectly.
Shanqi Liu, Wenzhou Chen, Guanzhong Tian, Jun Chen 0023, Yong Liu 0007
IEEE Trans. Neural Networks Learn. Syst.5
2023 Unified Data-Free Compression: Pruning and Quantization without Fine-Tuning
abstract
Structured pruning and quantization are promising approaches for reducing the inference time and memory footprint of neural networks. However, most existing methods require the original training dataset to fine-tune the model. This not only brings heavy resource consumption but also is not possible for applications with sensitive or proprietary data due to privacy and security concerns. Therefore, a few data-free methods are proposed to address this problem, but they perform data-free pruning and quantization separately, which does not explore the complementarity of pruning and quantization. In this paper, we propose a novel framework named Unified Data-Free Compression(UDFC), which performs pruning and quantization simultaneously without any data and fine-tuning process. Specifically, UDFC starts with the assumption that the partial information of a damaged(e.g., pruned or quantized) channel can be preserved by a linear combination of other channels, and then derives the reconstruction form from the assumption to restore the information loss due to compression. Finally, we formulate the reconstruction error between the original network and its compressed network, and theoretically deduce the closed-form solution. We evaluate the UDFC on the large-scale image classification task and obtain significant improvements over various network architectures and compression methods. For example, we achieve a 20.54% accuracy improvement on ImageNet dataset compared to SOTA method with 30% pruning ratio and 6-bit quantization on ResNet-34. Code will be available at here.
Shipeng Bai, Jun Chen 0023, Xintian Shen, Yixuan Qian, Yong Liu 0007
ICCV2
2023 Learning Global-aware Kernel for Image Harmonization
abstract
Image harmonization aims to solve the visual inconsistency problem in composited images by adaptively adjusting the foreground pixels with the background as references. Existing methods employ local color transformation or region matching between foreground and background, which neglects powerful proximity prior and independently distinguishes fore-/back-ground as a whole part for harmonization. As a result, they still show a limited performance across varied foreground objects and scenes. To address this issue, we propose a novel Global-aware Kernel Net-work (GKNet) to harmonize local regions with comprehensive consideration of long-distance background references. Specifically, GKNet includes two parts, i.e., harmony kernel prediction and harmony kernel modulation branches. The former includes a Long-distance Reference Extractor (LRE) to obtain long-distance context and Kernel Prediction Blocks (KPB) to predict multi-level harmony kernels by fusing global information with local features. To achieve this goal, a novel Selective Correlation Fusion (SCF) module is proposed to better select relevant long-distance background references for local harmonization. The latter employs the predicted kernels to harmonize foreground regions with local and global awareness. Abundant experiments demonstrate the superiority of our method for image harmonization over state-of-the-art methods, e.g., achieving 39.53dB PSNR that surpasses the best counterpart by +0.78dB ↑; decreasing fMSE/MSE by 11.5%↓/6.7%↓ compared with the SoTA method. Code will be available at here.
Xintian Shen, Jiangning Zhang, Jun Chen 0023, Shipeng Bai, Yabiao Wang, Chengjie Wang 0001, Yong Liu 0007
ICCV3
2023 SUBP: Soft Uniform Block Pruning for 1×N Sparse CNNs Multithreading Acceleration
Jingyang Xiang, Siqi Li 0009, Jun Chen 0023, Guang Dai, Shipeng Bai, Yukai Ma, Yong Liu 0007
NeurIPS3
2023 Single-shot pruning and quantization for hardware-friendly neural network acceleration
Bofeng Jiang, Jun Chen 0023, Yong Liu 0007
Eng. Appl. Artif. Intell.2
2023 Learning SpatioTemporal and Motion Features in a Unified 2D Network for Action Recognition
abstract
Recent methods for action recognition always apply 3D Convolutional Neural Networks (CNNs) to extract spatiotemporal features and introduce optical flows to present motion features. Although achieving state-of-the-art performance, they are expensive in both time and space. In this paper, we propose to represent both two kinds of features in a unified 2D CNN without any 3D convolution or optical flows calculation. In particular, we first design a channel-wise spatiotemporal module to present the spatiotemporal features and a channel-wise motion module to encode feature-level motion features efficiently. Besides, we provide a distinctive illustration of the two modules from the frequency domain by interpreting them as advanced and learnable versions of frequency components. Second, we combine these two modules and an identity mapping path into one united block that can easily replace the original residual block in the ResNet architecture, forming a simple yet effective network dubbed STM network by introducing very limited extra computation cost and parameters. Third, we propose a novel Twins Training framework for action recognition by incorporating a correlation loss to optimize the inter-class and intra-class correlation and a siamese structure to fully stretch the training data. We extensively validate the proposed STM on both temporal-related datasets (i.e., Something-Something v1 & v2) and scene-related datasets (i.e., Kinetics-400, UCF-101, and HMDB-51). It achieves favorable results against state-of-the-art methods in all the datasets.
Mengmeng Wang 0005, Jiazheng Xing, Jing Su 0005, Jun Chen 0023, Yong Liu 0007
IEEE Trans. Pattern Anal. Mach. Intell.4
2023 Data-free quantization via mixed-precision compensation without fine-tuning
Jun Chen 0023, Shipeng Bai, Tianxin Huang, Mengmeng Wang 0005, Guanzhong Tian, Yong Liu 0007
Pattern Recognit.1
2023 Multi-Dimension Compression of Feed-Forward Network in Vision Transformers
Shipeng Bai, Jun Chen 0023, Yu Yang 0001, Yong Liu 0007
Pattern Recognit. Lett.2
2022 Learning to Train a Point Cloud Reconstruction Network Without Matching
Tianxin Huang, Xuemeng Yang, Jiangning Zhang, Jinhao Cui, Jun Chen 0023, Xiangrui Zhao, Yong Liu 0007
ECCV (1)6
2022 Resolution-Free Point Cloud Sampling Network with Data Distillation
Tianxin Huang, Jiangning Zhang, Jun Chen 0023, Yuang Liu, Yong Liu 0007
ECCV (2)3
2022 SuperLine3D: Self-supervised Line Segmentation and Description for LiDAR Point Cloud
Xiangrui Zhao, Sheng Yang 0007, Tianxin Huang, Jun Chen 0023, Mingyang Li 0001, Yong Liu 0007
ECCV (9)4
2022 Fast Point Cloud Sampling Network
Tianxin Huang, Jun Chen 0023, Jiangning Zhang, Yong Liu 0007
Pattern Recognit. Lett.2
2022 3QNet: 3D Point Cloud Geometry Quantization Compression Network
abstract
Since the development of 3D applications, the point cloud, as a spatial description easily acquired by sensors, has been widely used in multiple areas such as SLAM and 3D reconstruction. Point Cloud Compression (PCC) has also attracted more attention as a primary step before point cloud transferring and saving, where the geometry compression is an important component of PCC to compress the points geometrical structures. However, existing non-learning-based geometry compression methods are often limited by manually pre-defined compression rules. Though learning-based compression methods can significantly improve the algorithm performances by learning compression rules from data, they still have some defects. Voxel-based compression networks introduce precision errors due to the voxelized operations, while point-based methods may have relatively weak robustness and are mainly designed for sparse point clouds. In this work, we propose a novel learning-based point cloud compression framework named 3D Point Cloud Geometry Quantiation Compression Network (3QNet), which overcomes the robustness limitation of existing point-based methods and can handle dense points. By learning a codebook including common structural features from simple and sparse shapes, 3QNet can efficiently deal with multiple kinds of point clouds. According to experiments on object models, indoor scenes, and outdoor scans, 3QNet can achieve better compression performances than many representative methods.
Tianxin Huang, Jiangning Zhang, Jun Chen 0023, Zhonggan Ding, Ying Tai, Zhenyu Zhang 0005, Chengjie Wang 0001, Yong Liu 0007
ACM Trans. Graph.3
2021 Pruning by Training: A Novel Deep Neural Network Compression Framework for Image Processing
abstract
Filter pruning for a pre-trained convolutional neural network is most normally performed through human-made constraints or criteria such as norms, ranks, etc. Typically, the pruning pipeline comprises two-stage: first learn a sparse structure from the original model, then optimize the weights in the new prune model. One disadvantage of using human-made criteria to prune filters is that the design and selection of threshold criteria depend on complicated prior knowledge. Besides, the pruning process is less robust due to the impact of directly regularizing on filters. To address the problems mentioned, we propose an effective one-stage pruning framework: introducing a trainable collaborative layer to jointly prune and learn neural networks in one go. In our framework, we first add a binary collaborative layer for each original filter. Then, a new type of gradient estimator - asymptotic gradient estimator is first introduced to pass the gradient in the binary collaborative layer. Finally, we simultaneously learn the sparse structure and optimize the weights from the original model in the training process. Our evaluation results on typical benchmarks, CIFAR and ImageNet, demonstrate very promising results against other state-of-the-art filter pruning methods.
Guanzhong Tian, Jun Chen 0023, Xianfang Zeng, Yong Liu 0007
IEEE Signal Process. Lett.2
2021 A Learning Framework for n-Bit Quantized Neural Networks Toward FPGAs
abstract
The quantized neural network (QNN) is an efficient approach for network compression and can be widely used in the implementation of field-programmable gate arrays (FPGAs). This article proposes a novel learning framework for n -bit QNNs, whose weights are constrained to the power of two. To solve the gradient vanishing problem, we propose a reconstructed gradient function for QNNs in the back-propagation algorithm that can directly get the real gradient rather than estimating an approximate gradient of the expected loss. We also propose a novel QNN structure named n -BQ-NN, which uses shift operation to replace the multiply operation and is more suitable for the inference on FPGAs. Furthermore, we also design a shift vector processing element (SVPE) array to replace all 16-bit multiplications with SHIFT operations in convolution operation on FPGAs. We also carry out comparable experiments to evaluate our framework. The experimental results show that the quantized models of ResNet, DenseNet, and AlexNet through our learning framework can achieve almost the same accuracies with the original full-precision models. Moreover, when using our learning framework to train our n -BQ-NN from scratch, it can achieve state-of-the-art results compared with typical low-precision QNNs. Experiments on Xilinx ZCU102 platform show that our n -BQ-NN with our SVPE can execute 2.9 times faster than that with the vector processing element (VPE) in inference. As the SHIFT operation in our SVPE array will not consume digital signal processing (DSP) resources on FPGAs, the experiments have shown that the use of SVPE array also reduces average energy consumption to 68.7% of the VPE array with 16 bit.
Jun Chen 0023, Liang Liu 0007, Yong Liu 0007, Xianfang Zeng
IEEE Trans. Neural Networks Learn. Syst.1