Tao Yang 0031

dblp:67/1120-31 · DBLP profile ↗
← Back
31ranked-venue papers
8as first author
30since 2021 · last 2026
0000-0001-8588-9483ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 30 · 8 first-author · 29 since 2021Software engineering, systems software and programming languages · 4 · 2 first-author · 4 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2026 When Low-Rank Meets Mixed-Precision: Training-Free Joint Compression for Efficient LLM Inference
abstract
The rapid growth of Large Language Models (LLMs) raises significant challenges for deployment in resource-constrained environments. Existing compression approaches, such as low-rank decomposition and quantization, are typically applied independently, which limits their effectiveness and fails to exploit their complementarity. To address this issue, we present a training-free framework for joint compression that integrates low-rank decomposition with mixed-precision quantization. We formulate the allocation of layer-wise rank and bit-width as a combinatorial optimization problem, guided by an input-aware sensitivity metric to allocate resources where they yield the highest accuracy retention. We further develop a sample-aware low-rank decomposition scheme with theoretical guarantees, and introduce a unified difference matrix to mitigate the coupled errors from structural approximation and quantization. Extensive experiments on diverse LLM architectures and datasets demonstrate that our method achieves state-of-the-art compression, reducing model size to 20% of the original while preserving inference accuracy. The code is available at https://github.com/zzzzzjq0126/HALO.git
Fangxin Liu, Jinqi Zhu, Chenyang Guan, Tao Yang 0031, Li Jiang 0002, Haibing Guan
ASP-DAC5
2025 Irregular Sparsity-Enabled Search-in-Memory Engine for Accelerating Spiking Neural Networks
Fangxin Liu, Zongwu Wang, Ning Yang 0012, Haomin Li 0002, Tao Yang 0031, Haibing Guan, Li Jiang 0002
APPT5
2025 STAMP: Accelerating Second-Order DNN Training Via ReRAM-Based Processing-in-Memory Architecture
Yilong Zhao 0004, Fangxin Liu, Mingyu Gao 0001, Xiaoyao Liang, Qidong Tang, Chengyang Gu, Tao Yang 0031, Naifeng Jing, Li Jiang 0002
APPT7
2025 PLAIN: Leveraging High Internal Bandwidth in PIM for Accelerating Large Language Model Inference via Mixed-Precision Quantization
abstract
DRAM-based processing-in-memory (DRAM-PIM) has gained commercial prominence in recent years. However, its integration for deep learning acceleration, particularly for large language models (LLMs), poses inherent challenges. Existing DRAM-PIM systems are limited in computational capabilities, primarily supporting element-wise and general matrix-vector multiplication (GEMV) operations, which contribute only a small portion of the execution time in LLM workloads. As a result, current systems still require powerful host processors to manage compute-heavy operations.To address these challenges and expand the applicability of commodity DRAM-PIMs in accelerating LLMs, we introduce PLAIN, a novel software/hardware co-design framework for PIM-enabled systems. PLAIN leverages the distribution locality of parameters and the unique characteristics of PIM to achieve optimal trade-offs between inference cost and model quality. Our framework includes three key innovations: 1) firstly, we propose a novel quantization algorithm that determines the optimal precision of parameters within each layer, considering both algorithmic and hardware characteristics to optimize hardware mapping; 2) PLAIN strategically utilizes both GPUs and PIMs, leveraging the high internal memory bandwidth within HBM for attention layers and the powerful compute capability of conventional systems for fully connected (FC) layers; 3) PLAIN integrates a workload-aware dataflow scheduler that efficiently arranges complex computations and memory access for mixed-precision tensors, optimizing execution across different hardware components. Experiments show PLAIN outperforms the conventional GPU with the same memory parameters and the state-of-the-art PIM accelerator, achieving a 5.03× and 1.69× performance boost, with negligible model quality loss.
Fangxin Liu, Zongwu Wang, Yilong Zhao 0004, Tao Yang 0031, Li Jiang 0002, Haibing Guan
ICCAD5
2025 QUARK: Quantization-Enabled Circuit Sharing for Transformer Acceleration by Exploiting Common Patterns in Nonlinear Operations
abstract
Transformer-based models have revolutionized computer vision (CV) and natural language processing (NLP) by achieving state-of-the-art performance across a range of benchmarks. However, nonlinear operations in models significantly contribute to inference latency, presenting unique challenges for efficient hardware acceleration. To this end, we propose QUARK, a quantization-enabled FPGA acceleration framework that leverages common patterns in nonlinear operations to enable efficient circuit sharing, thereby reducing hardware resource requirements. QUARK targets all nonlinear operations within Transformer-based models, achieving high-performance approximation through a novel circuit-sharing design tailored to accelerate these operations. Our evaluation demonstrates that QUARK significantly reduces the computational overhead of nonlinear operators in mainstream Transformer architectures, achieving up to a 1.96× end-to-end speedup over GPU implementations. Moreover, QUARK lowers the hardware overhead of nonlinear modules by more than 50% compared to prior approaches, all while maintaining high model accuracy—and even substantially boosting accuracy under ultra-low-bit quantization.
Zhixiong Zhao, Haomin Li 0002, Fangxin Liu, Yuncheng Lu, Zongwu Wang, Tao Yang 0031, Li Jiang 0002, Haibing Guan
ICCAD6
2025 SpMMPlu-Pro: An Enhanced Compiler Plug-In for Efficient SpMM and Sparsity Propagation Algorithm
abstract
Sparse matrix-matrix multiplication (SpMM) is a fundamental operation widely used in deep neural networks (DNNs) and high-performance computing. Many compilation studies have optimized the kernel code of SpMM to achieve better performance gains. However, on the one hand, these efforts often focus solely on optimizing individual SpMM operations, without fully considering the influence of preceding and subsequent operators on SpMM. On the other hand, when dense regions in SpMM require accumulation to the same output location, these dense matrix multiplications must be executed sequentially, leading to significant overhead from atomic additions or thread synchronization. In this article, we propose a novel compiler plug-in for efficient SpMM, named SpMMPlu-Pro. SpMMPlu-Pro inherits the sparse intermediate representation (Sparse IR) and sparse pattern representation [meta-operation (meta-op)] as well as five optimization passes from SpMMPlu. To fully utilize the sparse properties, SpMMPlu-Pro implements a forward and backward cross-layer sparsity propagation algorithm, which propagates the sparsity of one layer to the front and back layers, fully releasing the potential of utilizing sparsity to accelerate neural network inference. To alleviate the inefficient accumulation of meta-ops caused by atomic addition or thread synchronization, we propose two complementary scheduling schemes: 1) the segmentation and grouping algorithm based on automatic search and 2) the atomic optimization method through the meta-op data flow graph restructure. We integrated SpMMPlu-Pro into MindSpore and tested its effectiveness and scalability on the NVIDIA V100 GPU and Huawei Ascend 910. The results show that SpMMPlu-Pro supports various sparsity patterns, achieving an average speedup of$4.10\times $on the V100 GPU and$4.35\times $on the Ascend 910 compared to the dense counterpart.
Shiyuan Huang 0004, Fangxin Liu, Tao Yang 0031, Zongwu Wang, Ning Yang 0012, Li Jiang 0002
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3
2024 PAAP-HD: PIM-Assisted Approximation for Efficient Hyper-Dimensional Computing
abstract
Hyper-Dimensional Computing (HDC) is a brain-inspired learning framework that is particularly suited to resource-limited edge devices. HDC operates in a high-parallel manner, encoding raw data into hyper-dimensional space, thus enabling efficient training and inference. However, the high dimensionality of data representation in HDC demands a substantial multiplication cost for calculating cosine similarity in high-precision HDC processes. While binarization of HDC can circumvent these multiplications, it often results in unsatisfactory accuracy. In this paper, we propose PAAP-HD, a novel approximation framework that is both accurate and hardware-friendly, designed to enhance the efficiency of HDC inference. Our framework employs a simple neural network as a universal approximator, which can be mapped to parallel Multiply-Accumulate (MAC) operations of the ReRAM-based PIM crossbar. Additionally, we introduce an algorithm to guide model switching, which aids in managing the approximation quality. This algorithm can be instantiated as a just-in-time predictor, seamlessly integrated into HDC to prescribe the appropriate mode for each sample. Our evaluation is conducted on data sets in four different fields, and the results show that PAAP-HD can bring an execution time speedup of 93.1$\times$ and improve energy efficiency by 41.5$\times$ energy with just <1% accuracy loss.
Fangxin Liu, Haomin Li 0002, Ning Yang 0012, Yichi Chen 0001, Zongwu Wang, Tao Yang 0031, Li Jiang 0002
ASPDAC6
2024 TEAS: Exploiting Spiking Activity for Temporal-wise Adaptive Spiking Neural Networks
abstract
Spiking neural networks (SNNs) are energy-efficient alternatives to commonly used deep artificial neural networks (ANNs). However, their sequential computation pattern over multiple time steps makes processing latency a significant hindrance to deployment. In existing SNNs deployed on time-driven hardware, all layers generate and receive spikes in a synchronized manner, forcing them to share the same time steps. This often leads to considerable time redundancy in the spike sequences and considerable repetitive processing. Motivated by the effectiveness of dynamic neural networks for boosting efficiency, we propose a temporal-wise adaptive SNN, namely TEAS, in which each layer is configured with independent number of time steps to fully exploit the potential of SNNs. Specifically, given an SNN, the number of time steps of each layer is configured according to its contribution to the final performance of the whole network. Then, we exploit the temporal transforming module to produce a dynamic policy that can adapt the temporal information dynamically during inference. The adaptive configuration generating process enables trade-offs between model complexity and accuracy. Through extensive experiments on challenging datasets, we demonstrate that TEAS significantly improves energy efficiency and processing latency while achieving comparable accuracy to state-of-the-art methods.
Fangxin Liu, Haomin Li 0002, Ning Yang 0012, Zongwu Wang, Tao Yang 0031, Li Jiang 0002
ASPDAC5
2024 Sava: A Spatial- and Value-Aware Accelerator for Point Cloud Transformer
abstract
Point Cloud Transformer is undergoing a rising trend in both industry and academia. It aligns traditional point cloud feature extraction methods with the latest transformer architecture and achieves remarkable performance. However, accelerators for traditional point cloud neural networks (PCNNs) and those solely for transformers fail to capture the characteristics of point cloud transformers, thus exhibiting poor performance. To address this challenge, we propose Sava, a co-designed accelerator that adopts a spatial- and value-aware hybrid pruning strategy for point cloud transformers. In terms of the spatial domain, we observe that points in regions of various densities exhibit different levels of importance. In the value space, a minor input contributes less to features, indicating lower importance. Considering both perspectives, we hybridize the information inherited from the spatial and value spaces to prune less significant values in attention, which converts data to sparse patterns and makes it readily accelerated. Furthermore, we adopt low-bit quantization to boost computations and apply varying quantization precisions across different network layers based on their sensitivity. In support of our algorithm, we propose an architecture that employs a configurable mixed-precision systolic array for various computing loads under diverse precisions. To address the workload imbalance of the unstructured sparse computations, we introduce a data rearrangement mechanism, which improves resource utilization while hiding latency. We evaluate our Sava on four point cloud transformer models and achieve notable accuracy and performance gains. In comparison with CPU, GPUs, and ASICs, our Sava offers 10.3×, 3.6×, 3.3×, 2.6×, 2.2× speedup, along with 20×, 8.8×, 6.9×, 3.2×, 2.4× energy savings on average.
Xueyuan Liu 0001, Zhuoran Song, Xing Li 0031, Tao Yang 0031, Fangxin Liu, Xiaoyao Liang
DATE5
2024 T-BUS: Taming Bipartite Unstructured Sparsity for Energy-Efficient DNN Acceleration
abstract
Exploiting sparsity is a key technique to reduce the computation and memory cost attributed to the ever-expanding size of DNN models. Prior sparse DNN accelerators largely exploit structured sparsity, offering limited benefits due to the need to maintain lower sparsity levels to preserve the accuracy of the original models. On the other hand, exploiting unstructured sparsity requires complicated index accesses for non-zeros value. While this approach provides algorithmic advantages, it intro-duces significant hardware overheads due to irregular, largely unpredictable sparsity patterns. As such, it is not hardware-efficient and hence only achieves sub-optimal sparsity-exploiting benefits. To fully unleash the potential of unstructured sparsity, this paper introduces T-BUS, an algorithm and hardware co-design framework for an Efficient Unstructured Sparsity Engine. At the algorithm level, T-BUS proposes a novel sparse encoding format and computation ordering mechanism, reducing computation and storage costs simultaneously. At the hardware level, T-BUS incorporates a specialized parallel lookup structure with a novel dataflow for efficient index-matching operations in bilateral unstructured sparsity computations. Together, these techniques provide a practical approach to harness the highest potential benefits from non-structured sparsity in both storage and computation, while mitigating the challenges associated with unstructured sparsity in hardware design. Compared to existing works, T-BUS achieves up to 85.8% energy saving and 4.72x speedup across workloads with diverse unstructured sparsity levels.
Ning Yang 0012, Fangxin Liu, Zongwu Wang, Zhiyan Song, Tao Yang 0031, Li Jiang 0002
ICCD5
2024 UM-PIM: DRAM-based PIM with Uniform & Shared Memory Space
abstract
DRAM-based Processing in Memory (PIM) addresses the “memory wall” problem by incorporating computing units (PIM units) into main memory devices for faster and wider local data access. However, critical challenges prevent PIM units from being compatible with existing CPU hosts. Memory interleaving and virtual memory limit the size of contiguous data visible to PIM units that constrains the granularity of PIM tasks. Fine-grained PIM tasks result in significant CPU-PIM offloading overhead, offsetting the speed-up of PIM. Existing PIM systems adopt drastic measures to ensure PIM task offloading efficiency, including isolating PIM memory space and turning off global memory interleaving. These interventions, however, decrease the CPU’s memory bandwidth and introduce extra data transfer, leading to an additional “system memory wall”. This new “wall” must be eliminated before fully embracing the PIM technology. In this work, we propose UM-PIM, a PIM system with interleaved CPU pages and non-interleaved PIM pages coexisting in a Uniform and Shared Memory space. UM-PIM enables zero-copy during PIM task offloading and maintains the CPU’s memory bandwidth while ensuring PIM offloading efficiency. Firstly, we propose a dual-track memory management mechanism consisting of independent page allocation and address translation for the two kinds of pages, respectively. Second, we design UM-PIM interface hardware on the DIMM (with PIMs) side to provide a dynamic address mapping for accelerating the data re-layout. Finally, we provide APIs to reduce PIM-to-PIM communication overhead by optimizing the CPU’s access to PIM pages in different communication modes. We compare UM-PIM with a CPU system and the current PIM systems. Results show negligible performance degradation for CPU workloads ($\lt 0.1 \%$) on UM-PIM, contrasting with the $25.8 \%$ degradation on the current PIM system with memory interleaving switched off. For PIM workloads partitioned to CPU and PIM units, UM-PIM can reduce the CPU time by $4.93 \times$, resulting in an end-to-end $1.96 \times$ speedup on average.
Yilong Zhao 0004, Mingyu Gao 0001, Fangxin Liu, Zongwu Wang, Jin Li 0002, He Xian, Tao Yang 0031, Naifeng Jing, Xiaoyao Liang, Li Jiang 0002
ISCA10
2024 Exploiting Temporal-Unrolled Parallelism for Energy-Efficient SNN Acceleration
abstract
Event-driven spiking neural networks (SNNs) have demonstrated significant potential for achieving high energy and area efficiency. However, existing SNN accelerators suffer from issues such as high latency and energy consumption due to serial accumulation-comparison operations. This is mainly because SNN neurons integrate spikes, accumulate membrane potential, and generate output spikes when the potential exceeds a threshold. To address this, one approach is to leverage the sparsity of SNN spikes to reduce the number of time steps. However, this method can result in imbalanced workloads among neurons and limit the utilization of processing elements (PEs). In this paper, we present SATO, a temporal-parallel SNN accelerator that enables parallel accumulation of membrane potential for all time steps. SATO adopts a two-stage pipeline methodology, effectively decoupling neuron computations. This not only maintains accuracy but also unveils opportunities for fine-grained parallelism. By dividing the neuron computation into distinct stages, SATO enables the concurrent execution of spike accumulation for each time step, leveraging the parallel processing capabilities of modern hardware architectures. This not only enhances the overall efficiency of the accelerator but also reduces latency by exploiting parallelism at a granular level. The architecture of SATO includes a novel binary adder-search tree for generating the output spike train, effectively decoupling the chronological dependence in the accumulation-comparison operation. Furthermore, SATO employs a bucket-sort-based method to evenly distribute compressed workloads to all PEs, maximizing data locality of input spike trains. Experimental results on various SNN models demonstrate that SATO outperforms the well-known accelerator, the 8-bit version of “Eyeriss” by$20.7\times$in terms of speedup and$6.0\times$energy-saving, on average. Compared to the state-of-the-art SNN accelerator “SpinalFlow”, SATO can also achieve$4.6\times$performance gain and$3.1\times$energy reduction on average, which is quite impressive for inference.
Fangxin Liu, Zongwu Wang, Wenbo Zhao 0005, Ning Yang 0012, Yongbiao Chen, Shiyuan Huang 0004, Haomin Li 0002, Tao Yang 0031, Songwen Pei, Xiaoyao Liang, Li Jiang 0002
IEEE Trans. Parallel Distributed Syst.8
2023 HyperAttack: An Efficient Attack Framework for HyperDimensional Computing
abstract
HyperDimensional Computing (HDC) is emerging as a lightweight computational model for robust and efficient learning on resource-constrained hardware. Since HDC often runs on edge devices, the security challenge of HDC is a pressing issue confronting all the practitioners. Meanwhile, the security challenge of HDC’s parameters stored in memory has not been well studied. In this work, we are the first to propose a novel HDC attack framework called HyperAttack, which can crush a robust HDC model (i.e., binary HDC) by maliciously flipping an extremely few amount of bits within its memory system (i.e., DRAM) that stores the associative memory. Since the bit-flip operation can be conducted by the well-known Row Hammer attack, HyperAttack maximizes the accuracy degradation with the minimum number of bit-flips by identifying the bits closely related to the classification accuracy of hyperdimensional vectors (stored in the associative memory as binary vectors) in HDC. The proposed HyperAttack is based on the concept of fuzzing, combining dimensional ranking and distributions of features in hypervectors to identify the bits to be flipped. Our evaluation shows that HyperAttack can successfully attack a binary HDC by flipping only 10% bits of hyperdimensional vectors to decrease top-1 accuracy from 90.9% to 10%, while randomly flipping merely degrades the accuracy by less than 2%.
Fangxin Liu, Haomin Li 0002, Yongbiao Chen, Tao Yang 0031, Li Jiang 0002
DAC4
2023 SpMMPlu: A Compiler Plug-in with Sparse IR for Efficient Sparse Matrix Multiplication
abstract
Sparsity is becoming arguably the most critical dimension to explore for efficiency and scalability as deep learning models grow significantly larger. Particularly, pruning is a common method to reduce redundant computations in attention-based and convolution-based models. The induced sparse matrix multiplication (SpMM) normally requires domain-specific hardware architecture (DSA) to eliminate unnecessary zero-valued computations. However, generating an optimal kernel code for SpMM on general-purpose and ISA-based spatial accelerators without changing the hardware architecture is still an open problem.In this paper, we propose a compiler plug-in named SpMMPlu, which can extend the representation and optimization ability for SpMM in current deep learning compiler frameworks that only support dense matrix multiplication. The key of SpMMPlu is a flexible intermediate representation— Sparse IR, representing the SpMM with various sparsity patterns based on meta-ops with a multi-level structure. Meta-op takes abstraction of the hardware intrinsic as its minimum granularity, and the powerful optimizers of existing NN compiler backends (e.g., Auto-schedule in TVM, AKG in MindSpore) can be easily reused for its computational scheduling and code generation. Moreover, we propose a two-step (segmentation & grouping) method to achieve an efficient Sparse IR for each sparsity pattern. Only three passes are added in SpMMPlu to provide an automatic solution for SpMM kernel code generation. We embed SpMMPlu into MindSpore and do experiments on NVIDIA V100 GPU and Huawei Ascend 910 to verify its effectiveness and scalability. The results show that with SpMMPlu, MindSpore can support various sparsity patterns and deliver a 1.93× (on V100 GPU) and 2.21× (on AScend 910) speedup averagely compared to the dense counterpart.
Tao Yang 0031, Yiyuan Zhou, Qidong Tang, Jieru Zhao, Li Jiang 0002
DAC1
2023 PIMPR: PIM-based Personalized Recommendation with Heterogeneous Memory Hierarchy
abstract
Deep learning-based personalized recommendation models (DLRMs) are dominating AI tasks in data centers. The performance bottleneck of typical DLRMs mainly lies in the memory-bounded embedding layers. Resistive Random Access Memory (ReRAM)-based Processing-in-memory (PIM) architecture is a natural fit for DLRMs thanks to its in-situ computation and high computational density. However, it remains two challenges before DLRMs fully embrace ReRAM-based PIM architectures: 1) The size of DLRM's embedding tables can reach tens of GBs, far beyond the memory capacity of typical ReRAM chips. 2) The irregular sparsity conveyed in the embedding layers is difficult to exploit in ReRAM crossbars architecture. In this paper, we present a PIM-based DLRM accelerator named PIMPR. PIMPR has a heterogeneous memory hierarchy-ReRAM crossbar-based PIM modules serve as the computing caches with high computing parallelism, while DIMM modules are able to hold the entire embedding table-leveraging the data locality of DLRM's embedding layers. Moreover, we propose a runtime strategy to skip the useless calculation induced by the sparsity and an offline strategy to balance the workload of each ReRAM crossbar. Compared to the state-of-the-art DLRM accelerator SPACE and TRiM, PIMPR achieves on average 2.02×and 1.79× speedup, 5.6 ×, and 5.1 × energy reduction, respectively.
Tao Yang 0031, Yilong Zhao 0004, Fangxin Liu, Zhezhi He, Li Jiang 0002
DATE1
2023 SoBS-X: Squeeze-Out Bit Sparsity for ReRAM-Crossbar-Based Neural Network Accelerator
abstract
Resistive random-access-memory (ReRAM) crossbar is a promising technique for deep neural network (DNN) accelerators, thanks to its in-memory and in-situ analog computing abilities for vector–matrix multiplication-and-accumulations (VMMs). However, it is challenging for crossbar architecture to exploit the sparsity in DNNs. It is inevitably complex and costly to exploit fine-grained sparsity due to the limitation of the tightly coupled crossbar structure. As a countermeasure, we develop a novel ReRAM-based DNN accelerator, named sparse-multiplication-engine (SME), based on a hardware and software co-design framework. First, we orchestrate the bit-sparse pattern to increase the density of bit-sparsity based on existing quantization methods. Such quantized weights can be nicely generated using the alternating direction method of multipliers (ADMM) optimization during the DNN fine-tuning, which can exactly enforce bit patterns in weights. Second, we propose a novel weight mapping mechanism to slice the bits of the weight across crossbars and splice the activation results in peripheral circuits. This mechanism can decouple the tightly coupled crossbar structure and cumulate the sparsity in the crossbar. Finally, a superior squeeze-out scheme empties the crossbars mapped with highly sparse nonzeros from the previous two steps. We design the SME architecture and discuss its use for other quantization methods and different ReRAM cell technologies. We further propose a workload grouping algorithm and a pipeline to achieve workload balance among crossbar-rows that concurrently execute multiply–accumulate operations to optimize the system latency. Putting all together, with the optimized model, compared with prior state-of-the-art designs, the SME shrinks the use of crossbars up to$8.7\times $and$2.1\times $using ResNet-50 and MobileNet-v2, respectively, and achieve average$3.1\times $speed up with no or little accuracy loss on ImageNet.
Fangxin Liu, Zongwu Wang, Yongbiao Chen, Zhezhi He, Tao Yang 0031, Xiaoyao Liang, Li Jiang 0002
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.5
2023 PASGCN: An ReRAM-Based PIM Design for GCN With Adaptively Sparsified Graphs
abstract
Graph convolutional network (GCN) is a promising but computing- and memory-intensive learning model. Processing-in-memory (PIM) architecture based on the resistive random access memory-based crossbar (ReRAM crossbar) is a natural fit for GCN inference. It can reduce the data movements and compute the vector-matrix multiplication (VMM) in analog. However, it requires an unbearable crossbar cost to leverage the massive parallelism exhibited in GCNs. First, this article explores the design space for GCN inference on ReRAM crossbars and presents the first PIM-based GCN accelerator named PIMGCN, PIMGCN employs dense data mapping and a search-execute architecture to take full advantage of the intravertex parallelisms with acceptable crossbars cost. Two scheduling strategies for PIMGCN to maximize the intervertex parallelisms and optimize the pipeline are proposed. The optimal scheduling is reduced to a maximum independent set problem, which is solved by a novel node-grouping algorithm. Second, this article explores the task-irrelevant information in the graphs and proposes an adaptively sparsified GCN network targeted for PIMGCN, which is named as ASparGCN. ASparGCN exploits a multilayer perceptron (MLP)-based edge predictor to get edge selection strategies for each GCN layer separately and adaptively in the training stage, and only inferences with the selected edges in the test stage. We design two regularization terms to guide the selection strategies to achieve architecture-friendly sparse graphs for PIMGCN. The overall algorithm-architecture co-design is named as PASGCN. Compared to the state-of-the-art software framework running on Intel Xeon CPU and NVIDIA RTX8000 GPU, PASGCN achieves an average of$16455\times $and$110.7\times $speedup and 8.0E$+ 06\times $and 6.67E$+ 03\times $energy reduction, respectively. Compared with the ASIC accelerator HyGCN (Yan et al., 2020), PASGCN achieves$326.31\times $speedup and$124.8\times $energy reduction.
Tao Yang 0031, Zhuoran Song, Yilong Zhao 0004, Jiaxi Zhang 0001, Fangxin Liu, Li Jiang 0002
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.1
2023 DTATrans: Leveraging Dynamic Token-Based Quantization With Accuracy Compensation Mechanism for Efficient Transformer Architecture
abstract
Models based on the attention mechanism, i.e., transformers, have shown extraordinary performance in natural language processing (NLP) tasks. However, their memory footprint, inference latency, and power consumption are still prohibitive for efficient inference at edge devices, even at data centers. To tackle this issue, we present an algorithm-architecture co-design named DTATrans. We find empirically that the tolerance to the noise varies from token to token in attention-based NLP models. This finding leads us to dynamically quantize different tokens with mixed levels of bits. Furthermore, we find that the overstrict quantization method causes a dilemma of the model accuracy and model compression ratio, which impels us to explore a method to compensate for the model accuracy when the compression ratio is high. Thus, in DTATrans, we design a compression framework that: 1) dynamically quantizes tokens while they are forwarded in the models; 2) jointly determines the ratio of each precision; and 3) compensate the model accuracy by exploiting lightweight computing on the 0-bit tokens. Moreover, due to the dynamic mixed-precision tokens caused by our framework, previous matrix-multiplication accelerators (e.g., systolic array) cannot effectively exploit the benefit of the compressed attention computation. We thus design our transformer accelerator with the variable-speed systolic array (VSSA) and propose an effective optimization strategy to alleviate the pipeline-stall problem in VSSA without hardware overhead. We conduct experiments with existing attention-based NLP models, including BERT and GPT-2 on various language tasks. Our results show that DTATrans outperforms the previous neural network accelerator Eyeriss by$16.04\times $in terms of speedup and$3.62\times $in terms of energy saving. Compared with the state-of-the-art attention accelerator SpAtten, our DTATrans achieves at least$3.62\times $speedup and$4.22\times $energy efficiency improvement.
Tao Yang 0031, Fangxin Liu, Yilong Zhao 0004, Zhezhi He, Li Jiang 0002
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.1
2023 A Point Cloud Video Recognition Acceleration Framework Based on Tempo-Spatial Information
abstract
In point cloud video recognition (PVR) tasks, deep neural networks (DNNs) have been widely adopted to enhance accuracy. However, real-time processing is hindered due to the increasing volume of points and frames that require processes. Point clouds represent 3D-shaped discrete objects using a multitude of points. Consequently, these points often exhibit an uneven distribution in the view space, resulting in strong spatial similarity within each point cloud frame. Taking advantage of this observation, this article introduces PRADA, aPoint CloudRecognitionAcceleration algorithm viaDynamicApproximation. PRADA approximates and eliminates the similar local pairs’ computations and recovers their results by copying dissimilar local pairs’ features for speedup with negligible accuracy loss. Furthermore, considering the slow changes in point cloud frames that lead to the high temporal similarity among points across multiple frames, we design PointV, aPointCloudVideo Recognition Acceleration algorithm, to minimize unnecessary computations of similar points in the temporal domain. Moreover, we propose the PRADA and PointV architectures to accelerate the PRADA and PointV algorithms. These two architectures can be integrated to gain higher performance improvement. Our experiments on a wide variety of datasets show that PRADA averagely achieves about$7\times$speedup over 1080TI GPU. In addition, the experimental results show that the PointV architecture and the integrated architecture can respectively achieve$11.7\times$and$13.9\times$performance improvement with acceptable accuracy compared to the 1080TI GPU.
Zhuoran Song, Wanzhen Liu, Tao Yang 0031, Fangxin Liu, Naifeng Jing, Xiaoyao Liang
IEEE Trans. Parallel Distributed Syst.3
2022 PIM-DH: ReRAM-based processing-in-memory architecture for deep hashing acceleration
abstract
Deep hashing has gained growing momentum in large-scale image retrieval. However, deep hashing is computation- and memory-intensive, which demands hardware acceleration. The unique process of hash sequence computation in deep hashing is non-trivial to accelerate due to the lack of an efficient compute primitive for Hamming distance calculation and ranking.
Fangxin Liu, Wenbo Zhao 0005, Yongbiao Chen, Zongwu Wang, Zhezhi He, Qidong Tang, Tao Yang 0031, Cheng Zhuo, Li Jiang 0002
DAC8
2022 SATO: spiking neural network acceleration via temporal-oriented dataflow and architecture
abstract
Event-driven spiking neural networks (SNNs) have shown great promise for being strikingly energy-efficient. SNN neurons integrate the spikes, accumulate the membrane potential, and fire output spike when the potential exceeds a threshold. Existing SNN accelerators, however, have to carry out such accumulation-comparison operation in serial. Repetitive spike generation at each time step not only increases latency as well as overall energy budget, but also incurs memory access overhead of fetching membrane potentials, both of which lessen the efficiency of SNN accelerators. Meanwhile, inherent highly sparse spikes of SNNs lead to imbalanced workloads among neurons that hurdle the utilization of processing elements (PEs).
Fangxin Liu, Wenbo Zhao 0005, Zongwu Wang, Yongbiao Chen, Tao Yang 0031, Zhezhi He, Xiaokang Yang 0001, Li Jiang 0002
DAC5
2022 DTQAtten: Leveraging Dynamic Token-based Quantization for Efficient Attention Architecture
abstract
Models based on the attention mechanism, i.e. transformers, have shown extraordinary performance in Natural Language Processing (NLP) tasks. However, their memory footprint, inference latency, and power consumption are still prohibitive for efficient inference at edge devices, even at data centers. To tackle this issue, we present an algorithm-architecture co-design with dynamic and mixed-precision quantization, DTQAtten. We present empirically that the tolerance to the noise varies from token to token in attention-based NLP models. This finding leads us to quantize different tokens with mixed levels of bits. Thus, we design a compression framework that (i) dynamically quantizes tokens while they are forwarded in the models and (ii) jointly determines the ratio of each precision. Moreover, due to the dynamic mixed-precision tokens caused by our framework, previous matrix-multiplication accelerators (e.g. systolic array) cannot effectively exploit the benefit of the compressed attention computation. We thus design our accelerator with the variable-speed systolic array (VSSA) and propose an effective optimization strategy to alleviate the pipeline-stall problem in VSSA without hardware overhead. We conduct experiments with existing attention-based NLP models, including BERT and GPT-2 on various language tasks. Our results show that DTQAtten outperforms the previous neural network accelerator Eyeriss by 13.12× in terms of speedup and 3.8× in terms of energy-saving. Compared with the state-of-the-art attention accelerator SpAtten, our DTQAtten achieves at least 2.65× speedup and 3.38× energy efficiency improvement.
Tao Yang 0031, Zhuoran Song, Yilong Zhao 0004, Fangxin Liu, Zongwu Wang, Zhezhi He, Li Jiang 0002
DATE1
2022 Randomize and Match: Exploiting Irregular Sparsity for Energy Efficient Processing in SNNs
abstract
Spiking Neural Networks (SNNs) have emerged as a promising alternative to traditional deep Artificial Neural Networks (ANNs) due to its power efficiency that stems from their sparse spike-based computation. However, the spike train naturally exhibits high yet unbounded sparsity. This irregularity makes hardware inefficient if deployed directly on existing sparse CNN accelerators that strictly limit the sparsity patterns. Mean-while, SNN inherently contains a large number of redundant connections among neurons, which can be further exploited to reduce the computational burden on model deployment. Therefore, exploiting sparsity is a key technique in accelerating SNN inference on edge devices.To this end, we advocate exploiting irregular sparsity in SNNs for both input spikes (dynamic) and synaptic weights (static) since sparse spikes are inherently distributed in a random pattern and irregular sparsity is more flexible than regular ones. Thus, we propose MISS, a fraMework that takes full advantage of Irregular Sparsity in the SNN through synergistic hardware and software co-design. In the software part, we employ the unstructured pruning on the synaptic weights, eliminating the redundancy in network structure to the greatest extent without affecting the model accuracy. For the hardware part, we also design a sparsity-stationary dataflow that keeps sparse weights stationary in the memory to avoid the decoding overhead. With this dataflow and the matching-based architecture, we can efficiently unify the dynamic and static irregular sparsity to support the neuron computation with a very low overhead. Extensive evaluation on a wide variety of SNNs demonstrates that MISS achieves an average of 36% (up to 57%) improvement in energy efficiency and 23% (up to 48%) speedup over the baseline SNN accelerators.
Fangxin Liu, Zongwu Wang, Wenbo Zhao 0005, Yongbiao Chen, Tao Yang 0031, Xiaokang Yang 0001, Li Jiang 0002
ICCD5
2022 IVQ: In-Memory Acceleration of DNN Inference Exploiting Varied Quantization
abstract
Weight quantization is well adapted to cope with the ever-growing complexity of the deep neural network (DNN) model. Diversified quantization schemes lead to diverse quantized bit width and formats of the weights, thereby, subject to different hardware implementations. Such variety prevents a general NPU to leverage different quantization schemes to gain performance and energy efficiency. More importantly, a trend of quantization diversity emerges that applies multiple quantization schemes to different fine-grained structures (e.g., a layer or a channel of weight) of a DNN. Therefore, a general architecture is desired to exploit varied quantization schemes. The crossbar-based processing-in-memory (PIM) architecture, a promising DNN accelerator, is well known for its highly efficient matrix-vector multiplication. However, PIM suffers from the inflexible intracrossbar data path because the weight is stationary on the crossbar and binds to the “add” operation along the bitline. Therefore, many nonuniform quantization methods must rollback the quantization before mapping the weights onto the crossbar. Counterintuitively, this article discovers a unique opportunity of the PIM architecture to exploit varied quantization schemes. We first transform the quantization diversity problem into a consistency problem by aligning the bit with the same magnitude along the same bitline of the crossbar. Consequently, such naive weight mapping causes many square hollows of idle PIM cells. We then propose a novel spatial mapping to exempt these “hollow” crossbar from the intercrossbar data path. To further squeeze the weights on fewer crossbars, we decouple the intracrossbar data path from the hardware bitline by a novel temporal scheduling, so that bits with different magnitudes can be placed on cells along the same bitline. Finally, the proposed IVQ includes a temporal pipeline to avoid the introduced stalling cycles, and a data flow with delicate control mechanisms for the new intra and intercrossbar data paths. Putting all together, IVQ achieves$19.7\times $,$10.7\times $,$4.7\times \sim 63.4\times $,$91.7\times $speedup, and$17.7\times $,$5.1\times $,$5.7\times \sim 68.1\times $,$541\times $energy savings over two PIM accelerators (ISAAC and CASCADE), two customized quantization accelerators (based on ASIC and FPGA), and NVIDIA RTX 2080 GPU, respectively.
Fangxin Liu, Wenbo Zhao 0005, Zongwu Wang, Yilong Zhao 0004, Tao Yang 0031, Yiran Chen 0001, Li Jiang 0002
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.5
2021 AdaptiveGCN: Efficient GCN Through Adaptively Sparsifying Graphs
abstract
Graph Convolutional Networks (GCNs) have become the prevailing approach to efficiently learn representations from graph-structured data. Current GCN models adopt a neighborhood aggregation mechanism based on two primary operations, aggregation and combination. The workload of these two processes is determined by the input graph structure, making the graph input the bottleneck of processing GCN. Meanwhile, a large amount of task-irrelevant information in the graphs would hurt the model generalization performance. This brings the opportunity of studying how to remove the redundancy in the graphs. In this paper, we aim to accelerate GCN models by removing the task-irrelevant edges in the graph. We present AdaptiveGCN, an efficient and supervised graph sparsification framework. AdaptiveGCN adopts an edge predictor module to get edge selection strategies by learning the downstream task feedback signals for each GCN layer separately and adaptively in the training stage, then only inference with the selected edges in the test stage to speed up the GCN computation. The experimental results indicate that AdaptiveGCN could yield 43% (on CPU) and 39% (on GPU) GCN model speed-up averagely with comparable model performance on public graph learning benchmarks.
Tao Yang 0031, Lun Du, Zhezhi He, Li Jiang 0002
CIKM2
2021 PIMGCN: A ReRAM-Based PIM Design for Graph Convolutional Network Acceleration
abstract
Graph Convolutional Network (GCN) is a promising but computing- and memory-intensive learning model. Processing-in-memory (PIM) architecture based on the ReRAM crossbar is a natural fit for GCN inference. It can reduce the data movements and compute the vector-matrix multiplication (VMM) in analog. However, it requires an unbearable crossbar cost to leverage the massive parallelism exhibited in GCNs. This paper explores the design space for GCN acceleration on ReRAM crossbars and presents the first PIM-based GCN accelerator named PIMGCN. PIMGCN employs dense data mapping and a search-execute architecture to take full advantage of the intra-vertex parallelisms with acceptable crossbars cost. We further propose two scheduling strategies for PIMGCN to maximize the inter-vertex parallelisms and optimize the pipeline. The optimal scheduling is reduced to a maximum independent set problem, which is solved by a novel node-grouping algorithm. Compared to the state-of-the-art software framework running on Intel Xeon CPU and NVIDIA RTX8000 GPU, PIMGCN achieves on average 11044× and 74.3× speedup, 6.13E+06× and 5.09E+03× energy reduction, respectively. Compared with ASIC accelerator HyGCN [1], PIMGCN achieves 219× speedup and 95.3× energy reduction.
Tao Yang 0031, Yibo Han, Yilong Zhao 0004, Fangxin Liu, Xiaoyao Liang, Zhezhi He, Li Jiang 0002
DAC1
2021 IM3A: Boosting Deep Neural Network Efficiency via In-Memory Addressing-Assisted Acceleration
abstract
Most existing RRAM-based designs require expensive analog-to-digital converters (ADCs) digital-to-analog converters (DACs) and excessively occupied crossbars to achieve efficient acceleration. To reduce the overhead of DACs, the existing solution is to split the input into a bit sequence, but the MAC operation that can be completed by one cycle is forced to multiple cycles to the energy-efficiency decrease. For ADCs, it generally partitions the weight into multiple cells, resulting in an excessive number of crossbars or frequent writes on account of insufficient number. To solve this problem, we propose IM3A, an In-Memory Addressing-Assisted Acceleration scheme IM3A decompose MAC operations into multiplication and accumulation, which are implemented separately through the content-addressable and multiply-accumulated capabilities of the crossbar. The energy-efficiency is improved by the CAM crossbar supporting the parallel search of very large numbers of data bits, and the RRAM crossbar selectively enabling the rows to be read based on the hit result of the CAM search. Therefore, only the possibility of operands involved in MAC is deployed on the crossbar. Experimental results show that IM3A applied on various networks achieves system energy-efficiency improvement by 1.7x ∼ 15.9x over two state-of-the-art crossbar accelerators: ISAAC and PIM-Prune.
Fangxin Liu, Wenbo Zhao 0005, Zongwu Wang, Tao Yang 0031, Li Jiang 0002
ACM Great Lakes Symposium on VLSI4
2021 SME: ReRAM-based Sparse-Multiplication-Engine to Squeeze-Out Bit Sparsity of Neural Network
abstract
Resistive Random-Access-Memory (ReRAM) cross-bar is a promising technique for deep neural network (DNN) accelerators, thanks to its in-memory and in-situ analog computing abilities for Vector-Matrix Multiplication-and-Accumulations (VMMs). However, it is challenging for crossbar architecture to exploit the sparsity in DNNs. It inevitably causes complex and costly control to exploit fine-grained sparsity due to the limitation of tightly-coupled crossbar structure.As the countermeasure, we develop a novel ReRAM-based DNN accelerator, named Sparse-Multiplication-Engine (SME), based on a hardware and software co-design framework. First, we orchestrate the bit-sparse pattern to increase the density of bit-sparsity based on existing quantization methods. Second, we propose a novel weight mapping mechanism to slice the bits of a weight across the crossbars and splice the activation results in peripheral circuits. This mechanism can decouple the tightly-coupled crossbar structure and cumulate the sparsity in the crossbar. Finally, a superior squeeze-out scheme empties the crossbars mapped with highly-sparse non-zeros from the previous two steps. We design the SME architecture and discuss its use for other quantization methods and different ReRAM cell technologies. Compared with prior state-of-the-art designs, the SME shrinks the use of crossbars up to 8.7× and 2.1× using ResNet-50 and MobileNet-v2, respectively, with ≤ 0.3% accuracy drop on ImageNet.
Fangxin Liu, Wenbo Zhao 0005, Zhezhi He, Zongwu Wang, Yilong Zhao 0004, Tao Yang 0031, Jingnai Feng, Xiaoyao Liang, Li Jiang 0002
ICCD6
2021 An FPGA-Based Neural Network Overlay for ADAS Supporting Multi-Model and Multi-Mode
abstract
Advanced Driver-Assistance Systems (ADAS) are complex systems consisting of many computer vision tasks including image classification, object detection and semantic segmentation. FPGA is a feasible solution for deep learning based computer vision accelerator due to its high performance and energy efficiency. However, design a high performance FPGA accelerator requires good understanding of basic hardware concepts and consumes a long compilation time. Overlays can alleviate the above problems by accelerating applications in a software via a hardware architecture and a compiler. In this paper, we propose an FPGA-based neural network overlay processor for ADAS. The overlay architecture contains almost all common computation layers for learning based ADAS. In addition, we design a compiler that can automatically compile the high-level description of neural networks from deep learning framework like Caffe and Tensorflow into FPGA configurable codes, which can be executed by our overlay architecture without reprogramming. Experiments show that our overlay can process learning tasks in ADAS with low latency and low memory usage.
Jiaxi Zhang 0001, Tao Yang 0031, Qingzheng Li, Guojie Luo, Jianping Shi
ISCAS2
2021 BISWSRBS: A Winograd-based CNN Accelerator with a Fine-grained Regular Sparsity Pattern and Mixed Precision Quantization
abstract
Field-programmable Gate Array (FPGA) is a high-performance computing platform for Convolution Neural Networks (CNNs) inference. Winograd algorithm, weight pruning, and quantization are widely adopted to reduce the storage and arithmetic overhead of CNNs on FPGAs. Recent studies strive to prune the weights in the Winograd domain, however, resulting in irregular sparse patterns and leading to low parallelism and reduced utilization of resources. Besides, there are few works to discuss a suitable quantization scheme for Winograd. In this article, we propose a regular sparse pruning pattern in the Winograd-based CNN, namely, Sub-row-balanced Sparsity (SRBS) pattern, to overcome the challenge of the irregular sparse pattern. Then, we develop a two-step hardware co-optimization approach to improve the model accuracy using the SRBS pattern. Based on the pruned model, we implement a mixed precision quantization to further reduce the computational complexity of bit operations. Finally, we design an FPGA accelerator that takes both the advantage of the SRBS pattern to eliminate low-parallelism computation and the irregular memory accesses, as well as the mixed precision quantization to get a layer-wise bit width. Experimental results on VGG16/VGG-nagadomi with CIFAR-10 and ResNet-18/34/50 with ImageNet show up to 11.8×/8.67× and 8.17×/8.31×/10.6× speedup, 12.74×/9.19× and 8.75×/8.81×/11.1× energy efficiency improvement, respectively, compared with the state-of-the-art dense Winograd accelerator [20] with negligible loss of model accuracy. We also show that our design has 4.11× speedup compared with the state-of-the-art sparse Winograd accelerator [19] on VGG16.
Tao Yang 0031, Zhezhi He, Tengchuan Kou, Qingzheng Li, Haibao Yu, Fangxin Liu, Yun Liang 0001, Li Jiang 0002
ACM Trans. Reconfigurable Technol. Syst.1
2020 A Winograd-Based CNN Accelerator with a Fine-Grained Regular Sparsity Pattern
abstract
Field-Programmable Gate Array (FPGA) is a high-performance computing platform for Convolution Neural Networks (CNNs) inference. Winograd transformation and weight pruning are widely adopted to reduce the storage and arithmetic overhead in matrix multiplication of CNN on FPGAs. Recent studies strive to prune the weights in the Winograd domain, however, resulting in irregular sparse patterns and leading to low parallelism and reduced utilization of resources. In this paper, we propose a regular sparse pruning pattern in the Winograd-based CNN, namely Sub-Row-Balanced Sparsity (SRBS) pattern, to overcome the above challenge. Then, we develop a 2-step hardware co-optimization approach to improve the model accuracy using the SRBS pattern. Finally, we design an FPGA accelerator that takes advantage of the SRBS pattern to eliminate low-parallelism computation and irregular memory accesses. Experimental results on VGG16 and Resnet-18 with CIFAR-10 and Imagenet show up to 4.4x and 3.06x speedup compared with the state-of-the-art dense Winograd accelerator and 52% (theoretical upper-bound is 72%) performance enhancement compared with the state-of-the-art sparse Winograd accelerator. The resulting sparsity ratio is 80% and 75% and the loss of model accuracy is negligible.
Tao Yang 0031, Yunkun Liao, Jianping Shi, Yun Liang 0001, Naifeng Jing, Li Jiang 0002
FPL1