EDBT 2026 Demo / reviewers in the wild / expert
Ning Yang 0012
dblp:67/1751-12
· DBLP profile ↗
25ranked-venue papers
3as first author
25since 2021 · last 2026
0009-0004-6964-8910ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 24 · 3 first-author · 24 since 2021Software engineering, systems software and programming languages · 6 · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | EARTH: An Efficient MoE Accelerator with Entropy-Aware Speculative Prefetch and Result ReuseabstractMixture-of-Experts (MoE) models significantly reduce computation in large language models by activating only a subset of experts per input token, but they introduce severe memory bottlenecks due to the large number of expert parameters. Existing offloading and prefetching strategies either incur accuracy loss, prohibitively high memory traffic, or high decoding overhead, limiting deployment on resource-constrained hardware. In this work, we present EARTH, a hardware–software co-design that addresses these challenges through three key innovations. First, we propose a dual-entropy encoding scheme that decomposes each expert into a high-information base and a delta component, enabling compact storage while preserving accuracy via adaptive precision management. Second, we introduce a delta-aware speculative prefetching and reuse mechanism that preloads base components of predicted experts and selectively fetches deltas, reusing previously computed delta patterns to reduce memory traffic and redundant computation. Third, we design a hardware accelerator that is co-designed to efficiently support this encoding and prefetching strategy, optimizing execution order, parallelism, and memory utilization. Across representative MoE workloads, EARTH reduces data movement overhead, improves prefetch efficiency, and achieves up to 2.10× speedup compared to state-of-the-art baselines, while maintaining high model accuracy. Fangxin Liu, Ning Yang 0012, Jingkui Yang, Zongwu Wang, Chenyang Guan, Yu Feng 0007, Li Jiang 0002, Haibing Guan |
ASPLOS (2) | 2 |
| 2026 | STEP: Adaptive Spatio-Temporal Expert Prefetching for Low-Latency and Memory-Efficient MoE Inference
Fangxin Liu, Ning Yang 0012, Zongwu Wang, Chenyang Guan, Haomin Li 0002, Yu Feng 0007, Liqiang Lu, Siran Yang, Jiamang Wang, Lin Qu, Li Jiang 0002, Haibing Guan |
ISCA | 2 |
| 2026 | Harmonia: A Unified Hierarchical Scheduling Framework for Sparse Matrix Multiplication
Jingkui Yang, Fangxin Liu, Ning Yang 0012, Chenyang Guan, Zongwu Wang, Mei Wen, Li Jiang 0002, Haibing Guan |
ISCA | 4 |
| 2025 | Irregular Sparsity-Enabled Search-in-Memory Engine for Accelerating Spiking Neural Networks
Fangxin Liu, Zongwu Wang, Ning Yang 0012, Haomin Li 0002, Tao Yang 0031, Haibing Guan, Li Jiang 0002 |
APPT | 3 |
| 2025 | NeuronQuant: Accurate and Efficient Post-Training Quantization for Spiking Neural NetworksabstractSpiking neural networks (SNNs) are an alternative computational paradigm to artificial neural networks (ANNs) that have attracted attention due to their event-driven execution mechanisms, enabling extremely low energy consumption. However, a significant challenge and opportunity in SNNs is to optimize memory and compute costs while maintaining accuracy, thereby further reducing energy consumption. Model quantization has been proposed as a promising technique to improve the running efficiency via the number of data bits reduction. Whereas, this technique has yet to be well studied in the neuromorphic computing domain. The underlying reason is that the behaviors of SNNs are quite different from those of ANNs, making 1) the accuracy of SNNs usually sensitive to data precision, and 2) the introduction of a temporal dimension to characterize neuronal dynamics. In this paper, we present NeuronQuant, an accurate and energy-efficient quantization framework to reduce the precision of neurons while maintaining accuracy. The key insight is to design a post-training quantization method guided by the activity of neurons, efficiently reducing the bit-width of parameters based on local relationships within neurons. Additionally, a budget-aware mixed bit-width allocation strategy for the total model size enables the adaptive growth and narrowing of precision in each layer, leading to a mixed-precision quantization scheme of the desired size. Extensive evaluations demonstrate that NeuronQuant can achieve a compressed SNN with 5.2 bits on average and 1.51× power consumption reduction with a superior model accuracy which is quite impressive for SNN. Code is released at https://github.com/shieldforever/NeuronQuant. Haomin Li 0002, Fangxin Liu, Zewen Sun, Zongwu Wang, Shiyuan Huang 0004, Ning Yang 0012, Li Jiang 0002 |
ASP-DAC | 6 |
| 2025 | BLOOM: Bit-Slice Framework for DNN Acceleration with Mixed-PrecisionabstractDeep neural networks (DNNs) have revolutionized numerous AI applications, but their vast model sizes and limited hardware resources present significant deployment challenges. Model quantization offers a promising solution to bridge the gap between DNN size and hardware capacity. While INT8 quantization has been widely used, recent research has pushed for even lower precision, such as INT4. However, the presence of outliers-values with unusually large magnitudes-limits the effectiveness of current quantization techniques. Previous compression-based acceleration methods that incorporate outlieraware encoding introduce complex logic. A critical issue we have identified is that serialization and deserialization dominate the encoding/decoding time in these compression workflows, leading to substantial performance penalties during workflow execution. To address this challenge, we introduce a novel computing approach and a compatible architecture design named “BLOOM”. BLOOM leverages the strengths of the “bit-slicing” method, effectively combining structured mixed-precision and bit-level sparsity with adaptive dataflow techniques. The key insight of BLOOM is that outliers require higher precision, while normal values can be processed at lower precision. By interleaving 4-bit values, we efficiently exploit the inherent sparsity in the highprecision components. As a result, the BLOOM-based accelerator outperforms the existing outlier-aware accelerators by an average $1.2 \sim 4.0 \times$ speedup and $24.6 \% \sim 71.3 \%$ energy reduction, respectively, without model accuracy loss. Fangxin Liu, Ning Yang 0012, Zongwu Wang, Xuanpeng Zhu, Haidong Yao, Xiankui Xiong, Li Jiang 0002, Haibing Guan |
DAC | 2 |
| 2025 | PISA: Efficient Precision-Slice Framework for LLMs with Adaptive Numerical TypeabstractLarge language models (LLMs) have transformed numerous AI applications, with on-device deployment becoming increasingly important for reducing cloud computing costs and protecting user privacy. However, the astronomical model size and limited hardware resources pose significant deployment challenges. Model quantization is a promising approach to mitigate this gap, but the presence of outliers in LLMs reduces its effectiveness. Previous efforts addressed this issue by employing compression-based encoding for mixed-precision quantization. These approaches struggle to balance model accuracy with hardware efficiency due to their value-wise outlier granularity and complex encoding/decoding hardware logic. To address this, we propose PISA (Precision-Slice Framework), an acceleration framework that exploits massive sparsity in the higher-order part of LLMs by splitting 16-bit values into a 4-bit/12-bit format. Crucially, PISA introduces an early bird mechanism that leverages the high-order 4-bit computation to predict the importance of the full calculation result. This mechanism enables efficient computational skips by continuing execution only for important computations and using preset values for less significant ones. This scheme can be efficiently integrated with existing hardware accelerators like systolic arrays without complex encoding/decoding. As a result, PISA outperforms state-of-the-art precision-aware accelerators, achieving a $1.3-4.3 \times$ performance boost and $14.3-66.7 \%$ greater energy efficiency, with minimal model accuracy loss. This approach enables more efficient ondevice LLM deployment, effectively balancing computational efficiency and model accuracy. Ning Yang 0012, Zongwu Wang, Qingxiao Sun, Liqiang Lu, Fangxin Liu |
DAC | 1 |
| 2025 | TAIL: Exploiting Temporal Asynchronous Execution for Efficient Spiking Neural Networks with Inter-Layer ParallelismabstractSpiking neural networks (SNNs) are an alternative computational paradigm to artificial neural networks (ANNs) that have attracted attention due to their event-driven execution mechanisms, enabling extremely low energy consumption. However, the existing SNN execution model, based on software simulation or synchronized hardware circuitry, is incompatible with the event-driven nature, thus resulting in poor performance and energy efficiency. The challenge arises from the fact that neuron computations across multiple time steps result in increased latency and energy consumption. To overcome this bottleneck and leverage the full potential of SNNs, we propose TAIL, a pioneering temporal asynchronous execution mechanism for SNNs driven by a comprehensive analysis of SNN computations. Additionally, we propose an efficient dataflow design to support SNN inference, enabling concurrent computation of various time steps across multiple layers for optimal Processing Element (PE) utilization. Our evaluations show that TAIL greatly improves the performance of SNN inference, achieving a 6.94× speedup and a 6.97× increase in energy efficiency on current SNN computing platforms. Haomin Li 0002, Fangxin Liu, Zongwu Wang, Dongxu Lyu, Shiyuan Huang 0004, Ning Yang 0012, Zhuoran Song, Li Jiang 0002 |
DATE | 6 |
| 2025 | OPS: Outlier-Aware Precision-Slice Framework for LLM AccelerationabstractLarge language models (LLMs) have transformed numerous AI applications, with on-device deployment becoming increasingly important for reducing cloud computing costs and protecting user privacy. However, the astronomical model size and limited hardware resources pose significant deployment challenges. Model quantization is a promising approach to mitigate this gap, but the presence of outliers in LLMs reduces its effectiveness. Previous efforts addressed this issue by employing compression-based encoding for mixed-precision quantization. These approaches struggle to balance model accuracy with hard-ware efficiency due to their value-wise outlier granularity and complex encoding/decoding hardware logic. To address this, we propose OPS (Outlier-aware Precision-Slicing), an acceleration framework that exploits massive sparsity in the higher-order part of LLMs by splitting 16-bit values into a 4-bit/12-bit format. Crucially, OPS introduces an early bird mechanism that leverages the high-order 4-bit computation to predict the importance of the full calculation result. This mechanism enables efficient computational skips by continuing execution only for important computations and using preset values for less significant ones. This scheme can be efficiently integrated with existing hardware accelerators like systolic arrays without complex encoding/decoding. As a result, OPS outperforms state-of-the-art outlier-aware accelerators, achieving a 1.3 − 4.3× performance boost with minimal model accuracy loss. This approach enables more efficient on-device LLM deployment, effectively balancing computational efficiency and model accuracy. Fangxin Liu, Ning Yang 0012, Zongwu Wang, Xuanpeng Zhu, Haidong Yao, Xiankui Xiong, Li Jiang 0002 |
DATE | 2 |
| 2025 | CROSS: Compiler-Driven Optimization of Sparse DNNs Using Sparse/Dense Computation KernelsabstractAs deep learning models continue to grow larger and more complex, exploiting sparsity is becoming one of the most critical areas for enhancing efficiency and scalability. Several methods for leveraging sparsity have been proposed to more effectively balance the trade-off between compression ratio and accuracy. While these methods offer algorithmic advantages, they also introduce significant hardware overhead due to index-based encoding and decoding. In this paper, we propose CROSS, an end-to-end compilation optimization technique to achieve sparse DNN acceleration using GPU computation kernels. The key insight behind CROSS is to exploit parameter distribution locality and reconcile the “sparse” DNN computation with the high-performance “dense” computation kernels. Specifically, we perform an in-depth analysis of sparse operations in mainstream DNN computing frameworks. We then decompose the sparse workload into multiple components to create highly efficient, specialized operators with different sparsity levels. Additionally, we introduce a novel sparse graph translation technique to facilitate computation kernel processing of the sparse workload. The resulting CROSS framework can accommodate various sparsity patterns and optimization techniques, delivering an average $2.03 \times$ speedup on inference latency compared to seven state-of-the-art solutions with smaller memory footprints across various models and datasets. Fangxin Liu, Shiyuan Huang 0004, Ning Yang 0012, Zongwu Wang, Haomin Li 0002, Li Jiang 0002 |
HPCA | 3 |
| 2025 | FATE: Boosting the Performance of Hyper-Dimensional Computing Intelligence with Flexible Numerical DAta TypEabstractHyper-Dimensional Computing (HDC) is a promising braininspired learning framework designed for efficient, hardwarefriendly computation.By utilizing highly parallel operations, HDC encodes raw data into a hyper-dimensional space, facilitating efficient training and inference processes.However, the high precision required for representing high-dimensional vectors presents challenges for implementing HDC on resource-constrained edge devices, mainly due to the significant computational cost of performing multiplication for cosine similarity calculations.On the other hand, binary HDC offers lower costs but sacrifices accuracy.This paper addresses these challenges by focusing on data quantization, a hardware-efficient compression technique that incorporates sparsity to effectively balance cost and accuracy on embedded FPGA.Unlike existing methods that employ the same quantization scheme for all dimensions, we propose a novel solution that applies different numerical data types to different dimensions of data representations.This approach is motivated by two factors: * Both authors contributed equally to the paper. Haomin Li 0002, Fangxin Liu, Yichi Chen 0001, Zongwu Wang, Shiyuan Huang 0004, Ning Yang 0012, Dongxu Lyu, Li Jiang 0002 |
ISCA | 6 |
| 2025 | ASTER: Adaptive Dynamic Layer-Skipping for Efficient Transformer Inference via Markov Decision ProcessabstractTransformer-based models have demonstrated remarkable performance in computer vision tasks. However, their increasing model size leads to substantial memory demands and higher latency, hindering practical deployment. This paper presents an adaptive dynamic layer-skipping framework based on Markov Decision Process, which determines optimal computational paths based on the current state of input samples. We introduce a Temporal Importance Difference Reward mechanism to address the credit assignment problem in layer-skipping decisions, and develop a knowledge distillation strategy using learnable cognitive tokens to compensate for information loss. Experiments on various models demonstrate that our method significantly reduces computational costs while maintaining accuracy, offering a practical solution for deploying high-performance Transformer models in resource-constrained environments. The code is available at https://github.com/wjjkhl/ASTER Fangxin Liu, Ning Yang 0012, Zongwu Wang, Junping Zhao, Li Jiang 0002, Haibing Guan |
ACM Multimedia | 3 |
| 2025 | Attack and Defense: Enhancing Robustness of Binary Hyper-Dimensional ComputingabstractHyper-Dimensional Computing (HDC) has emerged as a lightweight computational model, renowned for its robust and efficient learning capabilities, particularly suitable for resource-constrained hardware. As HDC often finds its application in edge devices, the associated security challenges pose a critical concern that cannot be ignored. In this work, we aim to quantitatively delve into the robustness of binary HDC, which is widely recognized for its robustness. Employing the bit-flip attack as our initial focal point, we meticulously devise both an attack mechanism and a corresponding defense mechanism. Our objective is to comprehensively explore the robustness of the binary hyper-dimensional computation model, aiming to gain a deeper understanding of its security vulnerabilities and potential defenses. Specifically, we introduce a novel attack framework for HDC, named HyperAttack, which is capable of compromising a robust binary HDC model by maliciously flipping a minimal number of bits within its memory system (specifically, the DRAM) that houses the associative memory. The bit-flip operation is executed through the well-known Row Hammer attack, and HyperAttack optimizes the accuracy degradation by pinpointing the most vulnerable bits in the hyper-dimensional vectors (represented as binary vectors within the associative memory) of the HDC model. The proposed HyperAttack framework is grounded in the principles of fuzziness, seamlessly integrating dimensional ranking and feature similarity analysis within hypervectors to precisely identify the bits to be flipped. Furthermore, we have developed a defense mechanism named HyperDefense, designed to bolster the robustness of binary hyper-dimensional computational models against bit-flip attacks. This defense scheme is tailored specifically for HDC models, providing a robust safeguard against potential threats. HyperDefense operates directly on the associative memory of HDC models, strengthening their defenses. By meticulously modifying selected bits, HyperDefense maintains a high level of accuracy close to the original model, even in the face of increased bit flip rates. This defense mechanism leverages redundant dimensions as backups for critical information. Through a thorough analysis of dimension importance, HyperDefense achieves superior robustness by gracefully sacrificing non-critical dimensions, thus ensuring the model’s robustness against potential attacks. Haomin Li 0002, Fangxin Liu, Zongwu Wang, Ning Yang 0012, Shiyuan Huang 0004, Xiaoyao Liang, Haibing Guan, Li Jiang 0002 |
ACM Trans. Archit. Code Optim. | 4 |
| 2025 | SpMMPlu-Pro: An Enhanced Compiler Plug-In for Efficient SpMM and Sparsity Propagation AlgorithmabstractSparse matrix-matrix multiplication (SpMM) is a fundamental operation widely used in deep neural networks (DNNs) and high-performance computing. Many compilation studies have optimized the kernel code of SpMM to achieve better performance gains. However, on the one hand, these efforts often focus solely on optimizing individual SpMM operations, without fully considering the influence of preceding and subsequent operators on SpMM. On the other hand, when dense regions in SpMM require accumulation to the same output location, these dense matrix multiplications must be executed sequentially, leading to significant overhead from atomic additions or thread synchronization. In this article, we propose a novel compiler plug-in for efficient SpMM, named SpMMPlu-Pro. SpMMPlu-Pro inherits the sparse intermediate representation (Sparse IR) and sparse pattern representation [meta-operation (meta-op)] as well as five optimization passes from SpMMPlu. To fully utilize the sparse properties, SpMMPlu-Pro implements a forward and backward cross-layer sparsity propagation algorithm, which propagates the sparsity of one layer to the front and back layers, fully releasing the potential of utilizing sparsity to accelerate neural network inference. To alleviate the inefficient accumulation of meta-ops caused by atomic addition or thread synchronization, we propose two complementary scheduling schemes: 1) the segmentation and grouping algorithm based on automatic search and 2) the atomic optimization method through the meta-op data flow graph restructure. We integrated SpMMPlu-Pro into MindSpore and tested its effectiveness and scalability on the NVIDIA V100 GPU and Huawei Ascend 910. The results show that SpMMPlu-Pro supports various sparsity patterns, achieving an average speedup of$4.10\times $on the V100 GPU and$4.35\times $on the Ascend 910 compared to the dense counterpart. Shiyuan Huang 0004, Fangxin Liu, Tao Yang 0031, Zongwu Wang, Ning Yang 0012, Li Jiang 0002 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 5 |
| 2025 | STCO: Enhancing Training Efficiency via Structured Sparse Tensor Compilation OptimizationabstractNetwork sparsification serves as an effective technique to accelerate Deep Neural Network (DNN) inference. However, existing sparsification techniques often rely on structured sparsity, which yields limited benefits. This is primarily due to the significant memory and computational overhead introduced by numerous sparse storage formats during address generation and gradient updates. Additionally, many of these solutions are tailored solely for the inference phase, neglecting the crucial training phase. In this article, we introduce STCO, a novel Sparse Tensor Compilation Optimization technique that significantly enhances training efficiency through structured sparse tensor compilation. Central to STCO is the Tensorization-aware Index Entity (TIE) format, which effectively represents structured sparse tensors by eliminating redundant indices and minimizing storage overhead. The TIE format plays a pivotal role in the Address-carry flow (AC flow) pass, which optimizes the data layout at the computational graph level. This pass leverages the TIE format to enhance the efficiency of tensor representations, enabling more compact and efficient sparse tensor storage. Meanwhile, a shape inference pass utilizes the AC flow to derive optimized tensor shapes, further refining the performance of sparse tensor operations. Moreover, the Address-Carry TIE Flow dynamically tracks nonzero addresses, extending the benefits of sparse optimization to both forward and backward propagation. This seamless integration into the training pipeline enables a smooth transition to sparse tensor compilation without significant modifications to existing codebases. To further boost training performance, we implement an operator-level AC flow optimization pass tailored for structured sparse tensors. This pass generates efficient addresses, ensuring minimal computational overhead during sparse tensor operations. The flexibility of STCO allows it to be efficiently integrated into various frameworks or compilers, providing a robust solution for enhancing training efficiency with structured sparse tensors. Experiments demonstrated that STCO achieved impressive speedups of 3.64×, 5.43×, 4.89×, and 3.91× when compared to state-of-the-art sparse formats on VGG16, ResNet-18, MobileNetV1, and MobileNetV2, respectively. These findings underscore the efficiency and superiority of our proposed approach in leveraging unstructured sparsity for DNN inference acceleration. Shiyuan Huang 0004, Fangxin Liu, Zongwu Wang, Ning Yang 0012, Haomin Li 0002, Li Jiang 0002 |
ACM Trans. Design Autom. Electr. Syst. | 5 |
| 2024 | PAAP-HD: PIM-Assisted Approximation for Efficient Hyper-Dimensional ComputingabstractHyper-Dimensional Computing (HDC) is a brain-inspired learning framework that is particularly suited to resource-limited edge devices. HDC operates in a high-parallel manner, encoding raw data into hyper-dimensional space, thus enabling efficient training and inference. However, the high dimensionality of data representation in HDC demands a substantial multiplication cost for calculating cosine similarity in high-precision HDC processes. While binarization of HDC can circumvent these multiplications, it often results in unsatisfactory accuracy. In this paper, we propose PAAP-HD, a novel approximation framework that is both accurate and hardware-friendly, designed to enhance the efficiency of HDC inference. Our framework employs a simple neural network as a universal approximator, which can be mapped to parallel Multiply-Accumulate (MAC) operations of the ReRAM-based PIM crossbar. Additionally, we introduce an algorithm to guide model switching, which aids in managing the approximation quality. This algorithm can be instantiated as a just-in-time predictor, seamlessly integrated into HDC to prescribe the appropriate mode for each sample. Our evaluation is conducted on data sets in four different fields, and the results show that PAAP-HD can bring an execution time speedup of 93.1$\times$ and improve energy efficiency by 41.5$\times$ energy with just <1% accuracy loss. Fangxin Liu, Haomin Li 0002, Ning Yang 0012, Yichi Chen 0001, Zongwu Wang, Tao Yang 0031, Li Jiang 0002 |
ASPDAC | 3 |
| 2024 | TEAS: Exploiting Spiking Activity for Temporal-wise Adaptive Spiking Neural NetworksabstractSpiking neural networks (SNNs) are energy-efficient alternatives to commonly used deep artificial neural networks (ANNs). However, their sequential computation pattern over multiple time steps makes processing latency a significant hindrance to deployment. In existing SNNs deployed on time-driven hardware, all layers generate and receive spikes in a synchronized manner, forcing them to share the same time steps. This often leads to considerable time redundancy in the spike sequences and considerable repetitive processing. Motivated by the effectiveness of dynamic neural networks for boosting efficiency, we propose a temporal-wise adaptive SNN, namely TEAS, in which each layer is configured with independent number of time steps to fully exploit the potential of SNNs. Specifically, given an SNN, the number of time steps of each layer is configured according to its contribution to the final performance of the whole network. Then, we exploit the temporal transforming module to produce a dynamic policy that can adapt the temporal information dynamically during inference. The adaptive configuration generating process enables trade-offs between model complexity and accuracy. Through extensive experiments on challenging datasets, we demonstrate that TEAS significantly improves energy efficiency and processing latency while achieving comparable accuracy to state-of-the-art methods. Fangxin Liu, Haomin Li 0002, Ning Yang 0012, Zongwu Wang, Tao Yang 0031, Li Jiang 0002 |
ASPDAC | 3 |
| 2024 | INSPIRE: Accelerating Deep Neural Networks via Hardware-friendly Index-Pair EncodingabstractDeep Neural Network (DNN) inference consumes significant computing resources and development efforts due to the growing model size. Quantization is a promising technique to reduce the computation and memory cost of DNNs. Most existing quantization methods rely on fixed-point integers or floating-point types, which require more bits to maintain model accuracy. In contrast, variable-length quantization, which combines high precision for values with significant magnitudes (i.e., outliers) and low precision for normal values, offers algorithmic advantages but introduces significant hardware overhead due to variable-length encoding and decoding. Also, existing quantization methods are less effective for both (dynamic) activations and (static) weights due to the presence of outliers. Fangxin Liu, Ning Yang 0012, Zhiyan Song, Zongwu Wang, Haomin Li 0002, Shiyuan Huang 0004, Zhuoran Song, Songwen Pei, Li Jiang 0002 |
DAC | 2 |
| 2024 | EOS: An Energy-Oriented Attack Framework for Spiking Neural NetworksabstractSpiking neural networks (SNNs) are emerging as energy-efficient alternatives to traditional artificial neural networks (ANNs). Their event-driven information processing significantly reduces computational demands while maintaining competitive performance. However, as SNNs are increasingly deployed in edge devices, security concerns have emerged. While significant research efforts have been dedicated to addressing the security vulnerabilities stemming from malicious input, often referred to as adversarial examples, the security of SNN parameters remains relatively unexplored. This work introduces a novel attack methodology for SNNs known as Energy-Oriented SNN attack (EOS). EOS is designed to increase the energy consumption of SNNs through the malicious manipulation of binary bits within their memory systems (i.e., DRAM), where neuronal information is stored. The key insight of EOS lies in the observation that energy consumption in SNN implementations is intricately linked to spiking activity. The bit-flip operation, the well-known Row Hammer technique, is employed in EOS. It achieves this by identifying the most robust neurons in the SNN based on the spiking activity, particularly those related to the firing threshold, which is stored as binary bits in memory. EOS employs a combination of spiking activity analysis and a progressive search strategy to pinpoint the target neurons for bit-flip attacks. The primary objective is to incrementally increase the energy consumption of the SNN while ensuring that accuracy remains intact. With the implementation of EOS, successful attacks on SNNs can lead to an average of 43% energy increase with no drop in accuracy. Ning Yang 0012, Fangxin Liu, Zongwu Wang, Haomin Li 0002, Zhuoran Song, Songwen Pei, Li Jiang 0002 |
DAC | 1 |
| 2024 | SPARK: Scalable and Precision-Aware Acceleration of Neural Networks via Efficient EncodingabstractDeep Neural Networks (DNNs) have demonstrated remarkable success; however, their increasing model size poses a challenge due to the widening gap between model size and hardware capacity. To address this, model compression techniques have been proposed, but existing compression methods struggle to effectively handle the significant parameter variations (activations and weights) within the model. Moreover, current variance-aware encoding solutions for compression introduce complex logic, leading to limited compression benefits and hardware efficiency. In this context, we present SPARK, a novel algorithm/architecture co-designed solution that utilizes variable-length data representation for local parameter value processing, offering low hardware overhead and high-performance gains. Our key insight is that the high-order part in quantized values are often sparse, allowing us to employ an identity bit to assign the appropriate encoding length, thereby eliminating redundant bit-length footprints. This reduction in data representation based on data characteristics enables a serialized structured data encoding scheme that seamlessly integrates with existing hardware accelerators, such as systolic arrays. We evaluate SPARK-based accelerators against some existing encoding-based accelerator, and our results demonstrate significant improvements. The SPARK-based accelerator achieves up to 4.65 × speedup and 74.7% energy reduction, while maintaining superior model accuracy. Fangxin Liu, Ning Yang 0012, Haomin Li 0002, Zongwu Wang, Zhuoran Song, Songwen Pei, Li Jiang 0002 |
HPCA | 2 |
| 2024 | HOLES: Boosting Large Language Models Efficiency with Hardware-Friendly Lossless EncodingabstractTransformer-based large language models (LLMs) have demonstrated remarkable success; however, their increasing model size poses a challenge due to the widening gap between model size and hardware capacity. To address this, model com-pression techniques have been proposed, but existing compression methods struggle to effectively handle the significant parameter variations (activations and weights) within the model. More-over, current outlier-aware encoding solutions for compression introduce complex logic, leading to limited compression benefits and hardware efficiency. In this context, we present HOLES, a novel algorithm/architecture co-designed solution that utilizes variable-length data representation and metadata for local pa-rameter value processing, offering low hardware overhead and high-performance gains. Our key insight is that only a tiny fraction of activations are outliers demanding high precision representation. This observation enables us to exploit the sparsity of high significant bits within data representation, allowing us to employ an identifier bit to assign the appropriate encoding length, thereby eliminating redundant bit-length footprints. This reduction in data representation based on data characteristics enables a serialized structured data encoding scheme that seam-lessly integrates with existing hardware accelerators, such as systolic arrays. We evaluate HOLES-based accelerators against some existing encoding-based accelerator, and our results demon-strate significant improvements. The HOLES-based accelerator achieves up to 3.98 × speedup and 70.0% energy reduction, while maintaining superior model accuracy. Fangxin Liu, Ning Yang 0012, Zhiyan Song, Zongwu Wang, Li Jiang 0002 |
ICCD | 2 |
| 2024 | T-BUS: Taming Bipartite Unstructured Sparsity for Energy-Efficient DNN AccelerationabstractExploiting sparsity is a key technique to reduce the computation and memory cost attributed to the ever-expanding size of DNN models. Prior sparse DNN accelerators largely exploit structured sparsity, offering limited benefits due to the need to maintain lower sparsity levels to preserve the accuracy of the original models. On the other hand, exploiting unstructured sparsity requires complicated index accesses for non-zeros value. While this approach provides algorithmic advantages, it intro-duces significant hardware overheads due to irregular, largely unpredictable sparsity patterns. As such, it is not hardware-efficient and hence only achieves sub-optimal sparsity-exploiting benefits. To fully unleash the potential of unstructured sparsity, this paper introduces T-BUS, an algorithm and hardware co-design framework for an Efficient Unstructured Sparsity Engine. At the algorithm level, T-BUS proposes a novel sparse encoding format and computation ordering mechanism, reducing computation and storage costs simultaneously. At the hardware level, T-BUS incorporates a specialized parallel lookup structure with a novel dataflow for efficient index-matching operations in bilateral unstructured sparsity computations. Together, these techniques provide a practical approach to harness the highest potential benefits from non-structured sparsity in both storage and computation, while mitigating the challenges associated with unstructured sparsity in hardware design. Compared to existing works, T-BUS achieves up to 85.8% energy saving and 4.72x speedup across workloads with diverse unstructured sparsity levels. Ning Yang 0012, Fangxin Liu, Zongwu Wang, Zhiyan Song, Tao Yang 0031, Li Jiang 0002 |
ICCD | 1 |
| 2024 | COMPASS: SRAM-Based Computing-in-Memory SNN Accelerator with Adaptive Spike SpeculationabstractBrain-inspired spiking neural networks (SNNs) are considered energy-efficient alternatives to conventional deep neural networks (DNNs). By adopting event-driven information processing, SNNs can significantly reduce the computational demands associated with DNNs, while still achieving comparable performance. However, current SNNs primarily prioritize high accuracy by constructing complex neuron models that generate sparse spikes. Unfortunately, this approach results in low energy efficiency and high latency, posing a significant challenge for deploying SNNs at the edge. Furthermore, the dominant computation in SNNs, which involves spike-wise Accumulate-Compare operations, is well-suited for Computing-in-Memory (CIM) architectures. However, exploiting high parallel processing and spike sparsity in CIM-based SNN accelerators is challenging due to the irregularity and time dependency of spikes. To address these limitations, the paper proposes COMPASS, a SRAM-based CIM architecture for efficient SNNs. We first introduce an efficient method to exploit irregular sparsity for both input spikes (explicit) and output spikes (implicit). This is achieved through a speculation mechanism that exploit dynamic spike patterns, enabling lean hardware for sparsity utilization. Additionally, the CIM architecture is carefully modified to facilitate dynamic spike pattern generation and exploitation with minimal overhead. Moreover, we design an adaptive dataflow with temporal spike representation tailored for input/output spikes, reducing memory footprint and enabling parallel execution. Comprehensive evaluation results demonstrate that COMPASS can achieve 26.7x end-to-end speedup over recent SNN accelerators hardware implementation with up to 386.7x less energy per inference. Zongwu Wang, Fangxin Liu, Ning Yang 0012, Shiyuan Huang 0004, Haomin Li 0002, Li Jiang 0002 |
MICRO | 3 |
| 2024 | Exploiting Temporal-Unrolled Parallelism for Energy-Efficient SNN AccelerationabstractEvent-driven spiking neural networks (SNNs) have demonstrated significant potential for achieving high energy and area efficiency. However, existing SNN accelerators suffer from issues such as high latency and energy consumption due to serial accumulation-comparison operations. This is mainly because SNN neurons integrate spikes, accumulate membrane potential, and generate output spikes when the potential exceeds a threshold. To address this, one approach is to leverage the sparsity of SNN spikes to reduce the number of time steps. However, this method can result in imbalanced workloads among neurons and limit the utilization of processing elements (PEs). In this paper, we present SATO, a temporal-parallel SNN accelerator that enables parallel accumulation of membrane potential for all time steps. SATO adopts a two-stage pipeline methodology, effectively decoupling neuron computations. This not only maintains accuracy but also unveils opportunities for fine-grained parallelism. By dividing the neuron computation into distinct stages, SATO enables the concurrent execution of spike accumulation for each time step, leveraging the parallel processing capabilities of modern hardware architectures. This not only enhances the overall efficiency of the accelerator but also reduces latency by exploiting parallelism at a granular level. The architecture of SATO includes a novel binary adder-search tree for generating the output spike train, effectively decoupling the chronological dependence in the accumulation-comparison operation. Furthermore, SATO employs a bucket-sort-based method to evenly distribute compressed workloads to all PEs, maximizing data locality of input spike trains. Experimental results on various SNN models demonstrate that SATO outperforms the well-known accelerator, the 8-bit version of “Eyeriss” by$20.7\times$in terms of speedup and$6.0\times$energy-saving, on average. Compared to the state-of-the-art SNN accelerator “SpinalFlow”, SATO can also achieve$4.6\times$performance gain and$3.1\times$energy reduction on average, which is quite impressive for inference. Fangxin Liu, Zongwu Wang, Wenbo Zhao 0005, Ning Yang 0012, Yongbiao Chen, Shiyuan Huang 0004, Haomin Li 0002, Tao Yang 0031, Songwen Pei, Xiaoyao Liang, Li Jiang 0002 |
IEEE Trans. Parallel Distributed Syst. | 4 |
| 2023 | PSQ: An Automatic Search Framework for Data-Free Quantization on PIM-based ArchitectureabstractCrossbar-based Process-In-Memory (PIM) architecture has been considered as a promising solution for Deep Neural Networks (DNNs) acceleration. Due to the ever increasing model size and computational budget of DNNs, model compression is a critical step for the deployment of DNNs. However, when deploying DNNs in PIM architectures, fine-grained quantization on DNN weight matrices is not easy due to the inflexible data path inside the crossbar.To this end, in this paper, we study the feasibility and efficiency of a novel fine-grained quantization scheme called PSQ for PIM-based design. The scheme tightly combines the search principle of quantization and the PIM architecture to provide smooth hardware-friendly quantization. We leverage the weight locality and the variety of weight distributions in different blocks to facilitate the fine-grained quantization process. Meanwhile, we propose a lightweight search framework to adaptively allocate the quantization parameters (e.g., scale, bitwidth, etc.). During the search process, suitable quantization parameters are assigned directly to each fine-grained block, keeping the weight distributions before and after quantization as close as possible, thus minimizing the quantization errors. Our evaluation shows that the proposed PSQ achieves 3.5× reduction in occupied crossbars while the accuracy loss is negligible. What’s more, PSQ can perform such a process in just a few seconds on a single CPU, without model retraining and expensive computation. Fangxin Liu, Ning Yang 0012, Li Jiang 0002 |
ICCD | 2 |