VLDB 2026 Research / reviewers in the wild / expert
Gang Li 0015
dblp:62/2655-15
· DBLP profile ↗
32ranked-venue papers
4as first author
25since 2021 · last 2026
0000-0001-7835-4739ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 23 · 4 first-author · 20 since 2021Software engineering, systems software and programming languages · 9 · 2 first-author · 7 since 2021Artificial intelligence and machine learning · 7 · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Boosting the Performance of Tree-Based Speculative Decoding of LLMs on FPGAsabstractAs an efficient alternative to autoregressive decoding, tree-based speculative decoding (SD) has been widely adopted to accelerate LLM inference on GPUs. However, due to the notable disparity in compute power and memory bandwidth, we observe that a specific target-draft model pair with a proper decoding configuration, despite demonstrating significant performance gains on GPUs, often fails to maintain its efficacy on FPGAs, and may even underperform the standard autoregressive decoding approachIn this paper, we propose an analytical framework to revive the performance of tree-based speculative decoding on FPGAs. We introduce effective performance, a roofline-based metric designed to: 1) assess whether a specific target-draft model pair can benefit from SD for the given FPGA platform, and 2) determine the optimal decoding configuration to achieve peak performance when SD is applicable. We also propose a prior-score-based search strategy to identify the optimal tree structure for a preset number of nodes, further enhancing the performance. We evaluate our method on AMD FPGA platforms using two state-of-the-art SD algorithms: LongSpec and EAGLE-3. Our approach demonstrates a speedup of 2.54-3.89× over autoregressive decoding. Tielong Liu, Gang Li 0015, Zitao Mo, Minnan Pei, Jian Cheng 0001 |
DATE | 2 |
| 2026 | APEX: Integer-only Non-linear Function Approximation for Efficient Cross-Modal Inference
Peihuan Ni, Zitao Mo, Tielong Liu, Hongli Wen, Minnan Pei, Junwen Si, Weifan Guan, Peisong Wang 0001, Qinghao Hu 0001, Gang Li 0015, Jian Cheng 0001 |
DATE | 11 |
| 2026 | Towards efficient and accurate spiking neural networks via adaptive bit allocationabstractMulti-bit spiking neural networks (SNNs) have recently become a heated research spot, pursuing energy-efficient and high-accurate AI. However, with more bits involved, the associated memory and computation demands escalate to the point where the performance improvements become disproportionate. Based on the insight that different layers demonstrate different importance and extra bits could be wasted and interfering, this paper presents an adaptive bit allocation strategy for direct-trained SNNs, achieving fine-grained layer-wise allocation of memory and computation resources. Thus, SNN's efficiency and accuracy can be improved. Specifically, we parametrize the temporal lengths and the bit widths of weights and spikes, and make them learnable and controllable through gradients. To address the challenges caused by changeable bit widths and temporal lengths, we propose the refined spiking neuron, which can handle different temporal lengths, enable the derivation of gradients for temporal lengths, and suit spike quantization better. In addition, we theoretically formulate the step-size mismatch problem of learnable bit widths, which may incur severe quantization errors to SNN, and accordingly propose the step-size renewal mechanism to alleviate this issue. Experiments on various datasets, including the static CIFAR and ImageNet datasets and the dynamic CIFAR-DVS and DVS-GESTURE datasets, demonstrate that our methods can reduce the overall memory and computation cost while achieving higher accuracy. Particularly, our SEWResNet-34 can achieve a 2.69 % accuracy gain and 4.16 × lower bit budgets over the advanced baseline work on ImageNet. This work is open-sourced at this link. Xingting Yao, Qinghao Hu 0001, Tielong Liu, Gang Li 0015, Peisong Wang 0001, Jian Cheng 0001 |
Neural Networks | 5 |
| 2026 | MATA: A Memory-Efficient Attention Accelerator for LLMs Exploiting Look-Back KV Cache PruningabstractTransformer-based Large Language Models (LLMs) have sparked a new wave of AI applications. However, their large computational complexity and memory footprint pose significant challenges for real-world deployment. Although dedicated transformer accelerators have been widely explored, we observe that they are unefficient for decoder-only LLMs that feature autoregressive computations with KV Cache. Our in-depth analysis reveals that DRAM accesses induced by the KV Cache dominate the overall attention process. To address this issue, we propose aMemory-efficientATtentionAccelerator (MATA) for LLMs through algorithm and hardware co-design. Specifically,at the algorithm level, to mitigate the overhead caused by the linear increase of KV Cache, we propose a post-training Look-Back pruning method. It dynamically discards unimportant tokens through a comprehensive scoring scheme, thereby restricting KV Cache to a constant volume.At the hardware level, to identify important tokens with low latency, we design a Slice Top-K (STK) engine that can complete top-k-based sorting withO(N) time complexity. Moreover, we present the Adaptive Dataflow, which adaptively performs different inference phases of LLMs, thus significantly enhancing the PE array utilization. On average, our MATA can achieve speedups of 3.56×, 2.23× and 2.04×, 1.56× energy savings over two state-of-the-art transformer accelerators SpAtten and FACT, respectively. Gang Li 0015, Tielong Liu, Zitao Mo, Xiaoyao Liang, Jian Cheng 0001 |
IEEE Trans. Computers | 2 |
| 2025 | SBQ: Exploiting Significant Bits for Efficient and Accurate Post-Training DNN QuantizationabstractPost-Training Quantization is an effective technique for deep neural network acceleration. However, as the bit-width decreases to 4 bits and below, PTQ faces significant challenges in preserving accuracy, especially for attention-based models like LLMs. The main issue lies in considerable clipping and rounding errors induced by the limited number of quantization levels and narrow data range in conventional low-precision quantization. In this paper, we present an efficient and accurate PTQ method that targets 4 bits and below through algorithm and architecture co-design. Our key idea is to dynamically extract a small portion of significant bit terms from high-precision operands to perform low-precision multiplications under the given computational budget. Specifically, we propose Significant-Bit Quantization (SBQ). It exploits a product-aware method to dynamically identify significant terms and an error-compensated computation scheme to minimize compute errors. We present a dedicated inference engine to unleash the power of SBQ. Experiments on CNNs, ViTs, and LLMs reveal that SBQ consistently outperforms prior PTQ methods under 2~4-bit quantization. We also compare the proposed inference engine with state-of-the-art bit-operation-based quantization architectures TQ and Sibia. Results show that SBQ can achieve the highest area and energy efficiency. Jiayao Ling, Gang Li 0015, Qinghao Hu 0001, Xiaolong Lin, Jian Cheng 0001, Xiaoyao Liang |
DATE | 2 |
| 2025 | Light-DiT: An Importance-Aware Dynamic Compression Framework for Diffusion Transformers
Gang Li 0015, Xuan Zhang 0001, Jiayao Ling, Xiaolong Lin, Zhuoran Song, Jian Cheng 0001, Xiaoyao Liang |
Euro-Par (2) | 2 |
| 2025 | GSArch: Breaking Memory Barriers in 3D Gaussian Splatting Training via Architectural Supportabstract3D Gaussian Splatting (3DGS) introduces a novel methodology for representing scenes with anisotropic 3D Gaussian primitives, achieving exceptional quality and rendering speed in neural scene representation (NSR). However, the insufficient training speed of 3DGS limits its applicability in tasks that require online learning to perceive dynamic environments, such as autonomous driving and embodied intelligence. Although recent work, GSCore, has introduced a specialized accelerator for the rendering process of 3DGS, it overlooks the time-consuming backward propagation during 3DGS training.In this paper, we propose GSArch, a hardware architecture designed to overcome memory barriers and boost the efficiency of 3DGS training. Through a thorough characterization of 3DGS training, we identify three root causes of inefficiency: redundant data loading from off-chip memory, time-consuming atomic write operations, and severe bank conflicts during on-chip buffer reading. To address these challenges, GSArch introduces three architectural innovations. First, acknowledging that Gaussians vary in shape and often span multiple pixels, with larger Gaussians causing more repetitive data loading, GSArch employs hybrid memory management. This approach categorizes Gaussians into ‘hot’ and ‘cold’ ones, storing hot Gaussians in a fast but small on-chip buffer to reduce redundant loading while minimizing hardware costs. Second, GSArch leverages the informativeness variability of Gaussians’ gradients to filter out low-contribution gradients, significantly reducing atomic operations. Lastly, a rearrangement unit is designed to pack conflicting memory read requests into non-conflicting bundles. Our evaluation results demonstrate that GSArch achieves up to $6.49 \times$ and $15.42 \times$ speedups compared to Nvidia A100 and Jetson AGX Xavier, respectively, with substantially lower energy consumption and negligible image quality loss. Houshu He, Gang Li 0015, Fangxin Liu, Li Jiang 0002, Xiaoyao Liang, Zhuoran Song |
HPCA | 2 |
| 2025 | LISLLM: Long Context Inference of Large Language Models with Short KV CacheabstractLarge Language Models (LLMs) have demonstrated remarkable performance across various language tasks. However, in the process of generating long texts, the linearly increasing keyvalue (KV) cache imposes a large volume of memory footprint on HBM, resulting in significant time and energy consumption during inference. By retaining only the initial and recent tokens, existing methods ignore the intermediate tokens and the effect of the distance between current tokens and those in the KV cache, which causes significant accuracy degradation and limits the potential for KV cache compression. To overcome the above problems, we propose LISLLM, a KV cache compression mechanism that takes both the difference in importance and the distance between tokens into consideration for efficient inference of LLMs in long-context settings. Compared to the state-of-theart method, our method achieves compression ratios of up to$1.91 \times$and speedup of$1.20 \times$with better accuracy. Tielong Liu, Gang Li 0015, Zitao Mo, Xingting Yao, Jian Cheng 0001 |
ICPADS | 2 |
| 2025 | HEAT: NPU-NDP HEterogeneous Architecture for Transformer-Empowered Graph Neural NetworksabstractTransformer-empowered Graph Neural Networks (TF-GNNs) are gaining significant attention in AI research because they leverage the front-end Transformer's ability to process textual data while also harnessing the back-end GNN's capacity to analyze graph structures.Typically, TF-GNNs follow the sequential execution mode, where the front-end Transformer first encodes vertex features, followed by subgraph sampling and subsequent processing by the back-end GNN.However, due to the massive computation workloads of Transformers and the irregular memory access patterns of GNNs, achieving efficient inference for TF-GNNs remains a challenge.Although architectures like FACT and MEGA have been proposed to separately accelerate the Transformer and GNN, they overlook the new opportunities arising from the coupling of the Transformer and GNN.To enable efficient TF-GNNs, we propose HEAT, a heterogeneous architecture with a Neural Processing Unit (NPU) and a DIMM-based Near-Data Processing (NDP).Such a heterogeneous architecture can utilize both the high computational power of NPU and the high internal bandwidth of NDP.To fully unleash the potential of the NPU-NDP architecture, HEAT makes the following three contributions: First, HEAT leverages graph topology to identify the importance of vertices and encodes their features in the Transformer using varying precision accordingly.Second, HEAT gives more flexibility to the execution granularity and execution order of * Zhuoran Song is the corresponding author. Zhuoran Song, Yicheng Zheng, Gang Li 0015, Naifeng Jing, Xiaoyao Liang, Haibing Guan |
MICRO | 5 |
| 2025 | GCC: A 3DGS Inference Architecture with Gaussian-Wise and Cross-Stage Conditional Processingabstract3D Gaussian Splatting (3DGS) has emerged as a leading neural rendering technique for high-fidelity view synthesis, prompting the development of dedicated 3DGS accelerators for resource-constrained platforms.The conventional decoupled preprocessing-rendering dataflow in existing accelerators has two major limitations: 1) a Minnan Pei, Gang Li 0015, Junwen Si, Zitao Mo, Peisong Wang 0001, Zhuoran Song, Xiaoyao Liang, Jian Cheng 0001 |
MICRO | 2 |
| 2025 | An Efficient Bit-Sparse DNN Accelerator Exploiting Adaptive Bit-Serial ComputationsabstractBit sparsity, an intrinsic attribute of binary representation, has been widely utilized in DNN inference acceleration. Despite the advantages in performance and energy efficiency demonstrated by existing bit-serial-based bit-sparse accelerators, they still face two notable limitations: 1) At the low-level bit-serial multiplier level, existing methods either statically select weight or activation as the serialized object during the design phase, or simply serialize both without considering the distribution of non-zero bits in different operands, thereby failing to achieve optimal performance; 2) At the high-level dataflow level, existing approaches do not eliminate zero values in data movement and computation, leading to considerable energy and latency overhead, as well as suboptimal PE utilization. In this work, we propose AdaS-Pro accelerator for fast and energy-efficient DNN inference. At the multiplier level, AdaSPro employs an adaptive bit-serial computation scheme, which dynamically serializes the input operand with fewer non-zero bits at runtime, thereby minimizing compute cycles. To further enhance performance, AdaS-Pro introduces an improved Booth encoding method to reduce the number of non-zero bits in each operand. At the dataflow level, AdaS-Pro employs a compressed format to eliminate zero values and proposes a bi-directional inner-join unit coupled with a ring-shaped scheduler to achieve efficient non-zero workload extraction and balancing. Experimental results show that AdaS-Pro outperforms existing state-of-theart bit-sparse accelerators, such as BitLet, BitX, and Laconic, with performance improvements of 4.03×, 6.78×, and 1.43×, respectively. Jiayao Ling, Gang Li 0015, Xiaolong Lin, Xing Li 0031, Jian Cheng 0001, Xiaoyao Liang |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2024 | FastGL: A GPU-Efficient Framework for Accelerating Sampling-Based GNN Training at Large ScaleabstractGraph Neural Networks (GNNs) have shown great superiority on non-Euclidean graph data, achieving ground-breaking performance on various graph-related tasks. As a practical solution to train GNN on large graphs with billions of nodes and edges, the sampling-based training is widely adopted by existing training frameworks. However, through an in-depth analysis, we observe that the efficiency of existing sampling-based training frameworks is still limited due to the key bottlenecks lying in all three phases of sampling-based training, i.e., subgraph sample, memory IO, and computation. To this end, we propose FastGL, a GPU-efficient Framework for accelerating sampling-based training of GNN at Large scale by simultaneously optimizing all above three phases, taking into account both GPU characteristics and graph structure. Specifically, by exploiting the inherent overlap within graph structures, FastGL develops the Match-Reorder strategy to reduce the data traffic, which accelerates the memory IO without incurring any GPU memory overhead. Additionally, FastGL leverages a Memory-Aware computation method, harnessing the GPU memory's hierarchical nature to mitigate irregular data access during computation. FastGL further incorporates the Fused-Map approach aimed at diminishing the synchronization overhead during sampling. Extensive experiments demonstrate that FastGL can achieve an average speedup of 11.8×, 2.2× and 1.5× over the state-of-the-art frameworks PyG, DGL, and GNNLab, respectively. Our code is available at https://github.com/a1bc2def6g/fastgl-ae. Peisong Wang 0001, Qinghao Hu 0001, Gang Li 0015, Xiaoyao Liang, Jian Cheng 0001 |
ASPLOS (4) | 4 |
| 2024 | FusionArch: A Fusion-Based Accelerator for Point-Based Point Cloud Neural NetworksabstractPoint-based Point Cloud Neural Networks (PCNNs) have attracted much attention for their higher accuracy than voxel-based and multi-view-based PCNNs. Nevertheless, the increasing scale of point cloud data poses a challenge for real-time processing. Numerous previous works focus on accelerating PCNN inference but only optimize specific stages, limiting their generality to different networks with diverse performance bottlenecks. In this paper, we take nearly all stages of PCNNs into account, and propose 3 orthogonal algorithms, including Fusion-FPS, Fusion-Computation, and Fusion-Aggregation. We introduce Fusion-FPS to alter the sequential execution flow by reducing the Farthest Point Sampling (FPS) across layers to once and organize all neighbor search stages in parallel. To exclude redundant feature computations of “Filling Points”, we propose Fusion-Computation, identifying the presence and locations of “Filling Points” and directly borrowing the nearest neighbor features for them. To eliminate redundant memory accesses caused by shared neighbors in aggregation, we present Fusion-Aggregation, which clusters nearby centroids and coalesces their replicated accesses. In support of our algorithms, we co-design FusionArch, an architecture that implements our strategies and further optimizes memory access via a Local Fusion-Aggregation Table (LFT). We evaluate FusionArch on both server-level and edge-level platforms on 5 PCNNs across 4 applications and show remarkable accuracy and performance gains. On average, FusionArch achieves$2.6\times,5.6\times, 13.0\times$speedup and$17\times, 22\times, 62.4\times$energy savings over PointAcc.Server, NVIDIA AIOO GPU and Intel Xeon CPU, respectively. Moreover, it outperforms PRADA, PointAcc.Edge, Mesorasi and GPU with speedups of$2.4\times, 2.9\times, 5.3\times, 5.5\times$, and energy savings of$4.4\times, 7.2\times, 12.4\times, 11.5\times$, respectively. Xueyuan Liu 0001, Zhuoran Song, Guohao Dai 0001, Gang Li 0015, Can Xiao, Dehui Kong, Xiaoyao Liang |
DATE | 4 |
| 2024 | MEGA: A Memory-Efficient GNN Accelerator Exploiting Degree-Aware Mixed-Precision QuantizationabstractGraph Neural Networks (GNNs) are becoming a promising technique in various domains due to their excellent capabilities in modeling non-Euclidean data. Although a spectrum of accelerators has been proposed to accelerate the inference of GNNs, our analysis demonstrates that the latency and energy consumption induced by DRAM access still significantly impedes the improvement of performance and energy efficiency. To address this issue, we propose a Memory - Efficient GNN Accelerator (MEGA) through algorithm and hardware co-design in this work. Specifically, at the algorithm level, through an in-depth analysis of the node property, we observe that the data-independent quantization in previous works is not optimal in terms of accuracy and memory efficiency. This motivates us to propose the Degree-Aware mixed-precision quantization method, in which a proper bitwidth is learned and allocated to a node according to its in-degree to compress GNNs as much as possible while maintaining accuracy. At the hardware level, we employ a heterogeneous architecture design in which the aggregation and combination phases are implemented separately with different dataflows. In order to boost the performance and energy efficiency, we also present an Adaptive-Package format to alleviate the storage overhead caused by the fine-grained bitwidth and diverse sparsity, and a Condense-Edge scheduling method to enhance the data locality and further alleviate the access irregularity induced by the extremely sparse adjacency matrix in the graph. We implement our MEGA accelerator in a 28nm technology node. Extensive experiments demonstrate that MEGA can achieve an average speedup of 38.3 ×, 7.1 ×, 4.0 ×, 3.6× and 47.6 ×, 7.2 ×, 5.4 ×, 4.5 × energy savings over four state-of-the-art GNN accelerators, HyGCN, GCNAX, GROW, and SGCN, respectively, while retaining task accuracy. Fanrong Li, Gang Li 0015, Zejian Liu, Zitao Mo, Qinghao Hu 0001, Xiaoyao Liang, Jian Cheng 0001 |
HPCA | 3 |
| 2024 | GNeRF: Accelerating Neural Radiance Fields Inference via Adaptive Sample GatingabstractNeRF is an emerging algorithm in computer graphics that has achieved state-of-the-art results in areas such as image rendering and 3D reconstruction. However, to compute the RGB of pixels in a view, NeRF executes MLP calculations on a huge number of sample points, resulting in significant computational complexity. To address this issue, we propose a simple and hardware-friendly NeRF algorithm (dubbed GNeRF) in this paper. GNeRF is designed based on the concept of "gating-by-decomposing". Specifically, It decomposes the original large MLP into two smaller branches. For each ray, GNeRF utilizes one branch to predict the important samples based on the direction information adaptively. The RGB calculations are then solely performed on these important samples using the other branch. Experimental results show that GNeRF can achieve comparable PSNR with only 3% FLOPS of the original NeRF. To showcase the hardware efficiency of GNeRF, we also design an FPGA-based NeRF accelerator on Xilinx ZCU102 MPSoC. Evaluation reveals that GNeRF can significantly enhance inference performance with minimal modifications to the existing MLP engine. Gang Li 0015, Xiaolong Lin, Jiayao Ling, Xiaoyao Liang |
ISCAS | 2 |
| 2024 | Janus: A Flexible Processing-in-Memory Graph Accelerator Toward SparsityabstractGraph application is ever-growing in relational data analysis. However, the memory access patterns become the performance bottleneck in graph analytics and graph neural network (GNN) suffering from single-side and dual-side sparsity, separately. Existing resistive random access memory (RRAM)-based processing-in-memory accelerators reduce data movements but fail to handle both types of sparsity in graph data. To address these issues, our work introduces Janus, a flexible highly compact architecture that is capable of being configured to enable single-sparse mode and dual-sparse mode, to accelerate graph analytics and GNN workloads in compressed mapping, respectively. Upon performing graph analytics with single-side sparsity, Janus employs a tandem-isomorphic-crossbar design both to remove zero-stored footprint, and to eliminate redundant search and sequential indexing. To address the challenge of dual-side sparsity in GNN, Janus still takes a random index access mechanism to gather data rapidly and uses a semi-SPM2 compute paradigm to boost the RRAM-based analog multiplication-and-accumulation in the compressed format. Compared with the state-of-the-art works, Janus outperforms them in both performance and energy efficiency for graph analytics and GNN, respectively. Xing Li 0031, Zhuoran Song, Rachata Ausavarungnirun, Xiao Liu 0033, Xueyuan Liu 0001, Xuan Zhang 0001, Xuhang Wang, Jiayao Ling, Gang Li 0015, Naifeng Jing, Xiaoyao Liang |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 9 |
| 2023 | AdaS: A Fast and Energy-Efficient CNN Accelerator Exploiting Bit-SparsityabstractBit-sparsity has shown its promise in CNN acceleration. However, prior bit-sparse accelerators have two drawbacks: 1) a large number of zero values are involved in the computation and data movement; 2) the distribution of non-zero bits is not considered in PE design. To address these issues, we propose AdaS. At the multiplier level, we dynamically serialize the operands that have fewer non-zero bits. At the dataflow level, we propose a group-wise bi-directional inner-join for workload extraction and balancing. Results show that AdaS can achieve 3.28×, 2.05× speedup, and 1.99×, 1.80× energy efficiency over Bit-Pragmatic and Laconic, respectively. Xiaolong Lin, Gang Li 0015, Zizhao Liu, Zhuoran Song, Naifeng Jing, Xiaoyao Liang |
DAC | 2 |
| 2023 | PRADA: Point Cloud Recognition Acceleration via Dynamic ApproximationabstractRecent point cloud recognition (PCR) tasks tend to utilize deep neural network (DNN) for better accuracy. Still, the computational intensity of DNN makes them far from real-time processing, given the fast-increasing number of points that need to be processed. Because the point cloud represents 3D-shaped discrete objects in the physical world using a mass of points, the points tend for an uneven distribution in the view space that exposes strong clustering possibility and local pairs' similarities. Based on this observation, this paper proposes PRADA, an algorithm-architecture co-design that can accelerate PCR while reserving its accuracy. We propose dynamic approximation, which can approximate and eliminate the similar local pairs' computations and recover their results by copying key local pairs' features for PCR speedup without losing accuracy. For accuracy good, we further propose an advanced re-clustering technique to maximize the similarity between local pairs. For performance good, we then propose a PRADA architecture that can be built on any conventional DNN accelerator to dynamically approximate the similarity and skip the redundant DNN computation with memory accesses at the same time. Our experiments on a wide variety of datasets show that PRADA averagely achieves 4.2×, 4.9×, 7.1×, and 12.2× speedup over Mesorasi, V100 GPU, 1080TI GPU, and Xeon CPU with negligible accuracy loss. Zhuoran Song, Gang Li 0015, Li Jiang 0002, Naifeng Jing, Xiaoyao Liang |
DATE | 3 |
| 2023 | $\rm A^2Q$: Aggregation-Aware Quantization for Graph Neural Networks
Fanrong Li, Zitao Mo, Qinghao Hu 0001, Gang Li 0015, Zejian Liu, Xiaoyao Liang, Jian Cheng 0001 |
ICLR | 5 |
| 2023 | Extremely Sparse Networks via Binary Augmented Pruning for Fast Image ClassificationabstractNetwork pruning and binarization have been demonstrated to be effective in neural network accelerator design for high speed and energy efficiency. However, most existing pruning approaches achieve a poor tradeoff between accuracy and efficiency, which on the other hand, has limited the progress of neural network accelerators. At the same time, binary networks are highly efficient, however, a large accuracy gap exists between binary networks and their full-precision counterparts. In this article, we investigate the merits of extremely sparse networks with binary connections for image classification through software-hardware codesign. More specifically, we first propose a binary augmented extremely pruning method that can achieve ~98% sparsity with small accuracy degradation. Then we design the hardware architecture based on the resulting sparse and binary networks, which extensively explores the benefits of extreme sparsity with negligible resource consumption introduced by binary branch. Experiments on large-scale ImageNet classification and field-programmable gate array (FPGA) demonstrate that the proposed software-hardware architecture can achieve a prominent tradeoff between accuracy and efficiency. Peisong Wang 0001, Fanrong Li, Gang Li 0015, Jian Cheng 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2022 | PalQuant: Accelerating High-Precision Networks on Low-Precision Accelerators
Qinghao Hu 0001, Gang Li 0015, Qiman Wu, Jian Cheng 0001 |
ECCV (11) | 2 |
| 2022 | Ristretto: An Atomized Processing Architecture for Sparsity-Condensed Stream Flow in CNNabstractLow-precision quantization and sparsity have been widely explored in CNN acceleration due to their effectiveness in reducing computational complexity and memory requirements. However, to support variable numerical precision and sparse computation, prior accelerators design flexible multipliers or sparse dataflow separately. A uniform solution that simultaneously exploits mixed-precision and dual-sided irregular sparsity for CNN acceleration is still lacking. Through an in-depth review of existing precision-scalable and sparse accelerators, we observe that a direct combination of low-level multipliers and high-level sparse dataflow from both sides is challenging due to their orthogonal design spaces. To this end, in this paper, we propose condensed streaming computation. By representing non-zero weights and activations as atomized streams, the low-level mixed-precision multiplication and high-level sparse convolution can be unified into a shared dataflow through hierarchical data reuse. Based on the condensed streaming computation, we propose Ristretto, an atomized architecture that exploits both mixed-precision and dual-sided irregular sparsity for CNN inference. We implement Ristretto in a 28nm technology node. Extensive evaluations show that Ristretto consistently outperforms three state-of-the-art CNN accelerators, including Bit Fusion, Laconic, and SparTen, in terms of performance and energy efficiency. Gang Li 0015, Zhuoran Song, Naifeng Jing, Jian Cheng 0001, Xiaoyao Liang |
MICRO | 1 |
| 2022 | Block Convolution: Toward Memory-Efficient Inference of Large-Scale CNNs on FPGAabstractDeep convolutional neural networks have achieved remarkable progress in recent years. However, the large volume of intermediate results generated during inference poses a significant challenge to the accelerator design for resource-constrained field-programmable gate array (FPGA). Due to the limited on-chip storage, partial results of intermediate layers are frequently transferred back and forth between on-chip memory and off-chip DRAM, leading to a nonnegligible increase in latency and energy consumption. In this article, we propose block convolution, a hardware-friendly, simple, yet efficient convolution operation that can completely avoid the off-chip transfer of intermediate feature maps at runtime. The fundamental idea of block convolution is to eliminate the dependency of feature map tiles in the spatial dimension when spatial tiling is used, which is realized by splitting a feature map into independent blocks so that convolution can be performed separately on individual blocks. We conduct extensive experiments to demonstrate the efficacy of the proposed block convolution on both the algorithm side and the hardware side. Specifically, we evaluate block convolution on: 1) VGG-16, ResNet-18, ResNet-50, and MobileNet-V1 for the ImageNet classification task; 2) SSD and FPN for the COCO object detection task; and 3) VDSR for the Set5 single-image superresolution task. Experimental results demonstrate that comparable or higher accuracy can be achieved with block convolution. We also showcase two CNN accelerators via algorithm/hardware co-design based on block convolution on memory-limited FPGAs, and evaluation shows that both accelerators substantially outperform the baseline without off-chip transfer of intermediate feature maps. Gang Li 0015, Zejian Liu, Fanrong Li, Jian Cheng 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2021 | Hardware Acceleration of Fully Quantized BERT for Efficient Natural Language ProcessingabstractBERT is the most recent Transformer-based model that achieves state-of-the-art performance in various NLP tasks. In this paper, we investigate the hardware acceleration of BERT on FPGA for edge computing. To tackle the issue of huge computational complexity and memory footprint, we propose to fully quantize the BERT (FQ-BERT), including weights, activations, softmax, layer normalization, and all the intermediate results. Experiments demonstrate that the FQ-BERT can achieve 7.94× compression for weights with negligible performance loss. We then propose an accelerator tailored for the FQ-BERT and evaluate on Xilinx ZCU102 and ZCU11 FPGA. It can achieve a performance-per-watt of 3.18 fps/W, which is 28.91× and 12.72× over Intel(R) Core(TM) i7-8700 CPU and NVIDIA K80 GPU, respectively. Zejian Liu, Gang Li 0015, Jian Cheng 0001 |
DATE | 2 |
| 2021 | Dynamic Dual Gating Neural NetworksabstractIn dynamic neural networks that adapt computations to different inputs, gating-based methods have demonstrated notable generality and applicability in trading-off the model complexity and accuracy. However, existing works only explore the redundancy from a single point of the network, limiting the performance. In this paper, we propose dual gating, a new dynamic computing method, to reduce the model complexity at run-time. For each convolutional block, dual gating identifies the informative features along two separate dimensions, spatial and channel. Specifically, the spatial gating module estimates which areas are essential, and the channel gating module predicts the salient channels that contribute more to the results. Then the computation of both unimportant regions and irrelevant channels can be skipped dynamically during inference. Extensive experiments on a variety of datasets demonstrate that our method can achieve higher accuracy under similar computing budgets compared with other dynamic execution methods. In particular, dynamic dual gating can provide 59.7% saving in computing of ResNet50 with 76.41% top-1 accuracy on ImageNet, which has advanced the state-of-the-art. Codes are available at https://github.com/lfr-0531/DGNet. Fanrong Li, Gang Li 0015, Jian Cheng 0001 |
ICCV | 2 |
| 2020 | Sparsity-Inducing Binarized Neural NetworksabstractBinarization of feature representation is critical for Binarized Neural Networks (BNNs). Currently, sign function is the commonly used method for feature binarization. Although it works well on small datasets, the performance on ImageNet remains unsatisfied. Previous methods mainly focus on minimizing quantization error, improving the training strategies and decomposing each convolution layer into several binary convolution modules. However, whether sign is the only option for binarization has been largely overlooked. In this work, we propose the Sparsity-inducing Binarized Neural Network (Si-BNN), to quantize the activations to be either 0 or +1, which introduces sparsity into binary representation. We further introduce trainable thresholds into the backward function of binarization to guide the gradient propagation. Our method dramatically outperforms current state-of-the-arts, lowering the performance gap between full-precision networks and BNNs on mainstream architectures, achieving the new state-of-the-art on binarized AlexNet (Top-1 50.5%), ResNet-18 (Top-1 59.7%), and VGG-Net (Top-1 63.2%). At inference time, Si-BNN still enjoys the high efficiency of exclusive-not-or (xnor) operations. Peisong Wang 0001, Gang Li 0015, Tianli Zhao, Jian Cheng 0001 |
AAAI | 3 |
| 2020 | Hardware Acceleration of CNN with One-Hot Quantization of Weights and ActivationsabstractIn this paper, we propose a novel one-hot representation for weights and activations in CNN model and demonstrate its benefits on hardware accelerator design. Specifically, rather than merely reducing the bitwidth, we quantize both weights and activations into n-bit integers that containing only one non-zero bit per value. In this way, the massive multiply and accumulates (MACs) are equivalent to additions of powers of two that can be efficiently calculated with histogram based computaitons. Experiments on the ImageNet classification task show that comparable accuracy can be obtained on our proposed One-Hot Networks (OHN) compared to conventional fixed-point networks. As case studies, we evaluate the efficacy of the one-hot data representation on two state-of-the-art CNN accelerators on FPGA, our preliminary results show that 50% and 68.5% resource saving can be achieved on DaDianNao and Laconic respectively. Besides, the one-hot optimized Laconic can further achieve an average speedup of 4.94× on AlexNet. Gang Li 0015, Peisong Wang 0001, Zejian Liu, Cong Leng, Jian Cheng 0001 |
DATE | 1 |
| 2020 | Ladder Pyramid Networks For Single Image Super-ResolutionabstractBenefiting from the powerful representation capability of convolutional neural networks, the performance of single image super-resolution (SISR) has been substantially improved in recent years. However, many current CNN-based methods are computation-intensive because of large-size intermediate feature maps and inefficient convolutions. To resolve these problems, we propose Ladder Pyramid Network (LPN) for single image super-resolution. Firstly, we use strided convolution to reduce the size of the intermediate feature maps and thus reducing computation burden. In order to better balance the effectiveness and efficiency, we propose Ladder Pyramid Module to gradually fuse hierarchical features to enhance performance. Secondly, lightweight convolution block similar to Inverted Residual Module of Mobilenet-v2 was introduced into SISR, with which we build the network backbone and ladder feature pyramid. Experimental results demonstrate that the proposed Ladder Pyramid Network can achieve comparable or better performance than previous lightweight networks while reducing the amount of computation. Zitao Mo, Gang Li 0015, Jian Chen 0001 |
ICIP | 3 |
| 2020 | FSA: A Fine-Grained Systolic Accelerator for Sparse CNNsabstractSparsity, as an intrinsic property of convolutional neural networks (CNNs), has been widely employed for hardware acceleration, and many customized accelerators tailored for sparse weights or activations have been proposed in these years. However, the irregular sparse patterns introduced by both weights and activations are much more challenging for efficient computation. For example, due to the issues of access contention, workload imbalance, and tile fragmentation, the state-of-the-art sparse accelerator SCNN fails to fully leverage the benefits of sparsity, leading to nonoptimal results for both speedup and energy efficiency. In this article, we propose an efficient sparse CNN accelerator for both weights and activations, namely finegrained systolic accelerator (FSA), which jointly optimizes both hardware dataflow and software partitioning and scheduling strategy. Specifically, to deal with the access contentions problem, we present a fine-grained systolic dataflow, in which the activations move rhythmically along the horizontal processing element array while the weights are fed into the array in a fine-grained order. We then propose a hybrid network partitioning strategy that sets different partitioning strategies for different layers to balance the workload and alleviate the fragmentation problem caused by both sparse weights and activations. Finally, we present a scheduling search strategy to find the optimized schedules for neural networks, which can further improve energy efficiency. Extensive evaluations show that the proposed FSA consistently outperforms SCNN over AlexNet, VGGNet, GoogLeNet, and ResNet with an average speedup of 1.74× and up to 13.86× energy efficiency. Fanrong Li, Gang Li 0015, Zitao Mo, Jian Cheng 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2018 | Block convolution: Towards memory-efficient inference of large-scale CNNs on FPGAabstractFPGA-based CNN accelerators are gaining popularity due to high energy efficiency and great flexibility in recent years. However, as the networks grow in depth and width, the great volume of intermediate data is too large to store on chip, data transfers between on-chip memory and off-chip memory should be frequently executed, which leads to unexpected offchip memory access latency and energy consumption. In this paper, we propose a block convolution approach, which is a memory-efficient, simple yet effective block-based convolution to completely avoid intermediate data from streaming out to off-chip memory during network inference. Experiments on the very large VGG-16 network show that the improved top-1/top-5 accuracy of 72.60%/91.10% can be achieved on the ImageNet classification task with the proposed approach. As a case study, we implement the VGG-16 network with block convolution on Xilinx Zynq ZC706 board, achieving a frame rate of 12.19fps under 150MHz working frequency, with all intermediate data staying on chip. Gang Li 0015, Fanrong Li, Tianli Zhao, Jian Cheng 0001 |
DATE | 1 |
| 2018 | Training Binary Weight Networks via Semi-Binary Decomposition
Qinghao Hu 0001, Gang Li 0015, Peisong Wang 0001, Yifan Zhang 0001, Jian Cheng 0001 |
ECCV (13) | 2 |
| 2018 | Recent advances in efficient computation of deep convolutional neural networksabstractDeep neural networks have evolved remarkably over the past few years and they are currently the fundamental tools of many intelligent systems. At the same time, the computational complexity and resource consumption of these networks continue to increase. This poses a significant challenge to the deployment of such networks, especially in real-time applications or on resource-limited devices. Thus, network acceleration has become a hot topic within the deep learning community. As for hardware implementation of deep neural networks, a batch of accelerators based on a field-programmable gate array (FPGA) or an application-specific integrated circuit (ASIC) have been proposed in recent years. In this paper, we provide a comprehensive survey of recent advances in network acceleration, compression, and accelerator design from both algorithm and hardware points of view. Specifically, we provide a thorough analysis of each of the following topics: network pruning, low-rank approximation, network quantization, teacher–student networks, compact network design, and hardware accelerators. Finally, we introduce and discuss a few possible future directions. Jian Cheng 0001, Peisong Wang 0001, Gang Li 0015, Qinghao Hu 0001, Hanqing Lu |
Frontiers Inf. Technol. Electron. Eng. | 3 |