VLDB 2026 Research / reviewers in the wild / expert
Ling Liang 0003
dblp:95/7436-3
· DBLP profile ↗
43ranked-venue papers
6as first author
38since 2021 · last 2026
0000-0002-8534-6494ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 31 · 3 first-author · 28 since 2021Artificial intelligence and machine learning · 11 · 2 first-author · 9 since 2021Software engineering, systems software and programming languages · 8 · 1 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | FlexMem: High-Parallel Near-Memory Architecture for Flexible Dataflow in Fully Homomorphic EncryptionabstractFully Homomorphic Encryption (FHE) imposes substantial memory demands, presenting significant challenges for efficient hardware acceleration. Near-Memory Processing (NMP) has emerged as a promising architectural solution to alleviate the memory bottleneck. However, the irregular memory access patterns and flexible dataflows inherent to FHE limit the effectiveness of existing NMP accelerators, which fail to fully utilize the available near-memory bandwidth. In this work, we propose FlexMem, a near-memory accelerator featuring high-parallel computational units with varying memory access strides and interconnect topologies to effectively handle irregular memory access patterns. Furthermore, we design polynomialand ciphertext-level dataflows to efficiently utilize near-memory bandwidth under varying degrees of polynomial parallelism and enhance parallel performance. Experimental results demonstrate that FlexMem achieves $1.26 \times$ performance improvement over the state-of-the-art near-memory architectures in end-to-end benchmarks, with on average 95.7% of near-memory bandwidth utilization. Shangyi Shi, Husheng Han, Jianan Mu, Xinyao Zheng, Ling Liang 0003, Zidong Du, Xiaowei Li 0001, Xing Hu 0001 |
ASP-DAC | 5 |
| 2026 | CHIP-MAP: A Collaborative Optimization Framework for Macro Placement Using Large Language ModelsabstractAs integrated circuits continue to grow in both scale and complexity, macro placement plays a critical role in physical design, directly affecting chip-level performance, power, and area (PPA). Traditional macro placement methods, such as simulated annealing, analytical optimization, and reinforcement learning, face limitations including slow convergence, heavy dependence on large datasets, and over-reliance on intermediate PPA indicators rather than final PPA. Large language models (LLMs) offer strong generative power and semantic reasoning that can potentially automate macro layout tasks while addressing the aforementioned problems in traditional methods, but their limited understanding of layout rules and lack of iterative, feedback-driven refinement make direct application challenging. To address this, we propose CHIP-MAP, a macro placement framework based on multi-agent collaboration and feedback-driven optimization. Furthermore, we introduce two innovative tools: the Module Link Weight Analyzer (MWA) and the Standard Cell Usability Score (SCUS), which are designed to guide fine-grained layout refinement. We evaluate CHIP-MAP on five benchmarks ranging from low-power cores to large multi-core processors implemented at 130nm and 45nm technology nodes. Results show that it achieves up to 1.5% area reduction and an average repair of 61.6% of total negative slack (TNS), while also reducing wirelength and improving timing. Yiming Du, Renye Yan, Yunfan Yang, Frank Qu, Jiajun Tan, ZhiYu Zheng, Yiming Gan, Ling Liang 0003, Zongwei Wang 0001, Yimao Cai |
DATE | 8 |
| 2026 | GMaC: NvCIM Architecture for Parallel Point-based Point Cloud Acceleration via Geometric Mapping and Address-Index Computation
Zongwei Wang 0001, Ling Liang 0003, Yimao Cai |
DATE | 3 |
| 2026 | He2: A Communication-Light Heterogeneous Architecture for Efficient Fully Homomorphic Encryption
Shangyi Shi, Husheng Han, Zhaoxuan Kan, Jianan Mu, Tenghui Hua, Xinyao Zheng, Ling Liang 0003, Zidong Du, Xing Hu 0001 |
ISCA | 9 |
| 2026 | A Near-Sensor Image Compression Architecture with RRAM-based Hyperdimensional Encoder for Smart Vision
Haoyang Gu, Zongwei Wang 0001, Jingshan Li, Zezhi Chen, Ling Liang 0003, Wengao Lu, Yimao Cai |
ISCAS | 6 |
| 2026 | Efficient LoRA-Based Weight Update Write-Back in 3D NAND Flash for Large Language Models
Dongxue Zhao, Tianyang Luo, Ling Liang 0003, Zongwei Wang 0001, Yimao Cai |
ISCAS | 3 |
| 2026 | eBrainISA: Edge-Oriented Instruction Set Architecture for Hybrid Brain-Inspired Computing
Yujie Ying, Ziyi Yang 0014, Ling Liang 0003, Zegang Peng, Yifan Hu 0013, Zhuo Zou, Lei Deng 0003 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 5 |
| 2026 | OmniGuard: Two-Level Protection Framework for RRAM-Based Accelerator With High Efficiency and FlexibilityabstractRRAM-based Deep Neural Network (DNN) accelerators have gained widespread usage in edge devices. However, the security vulnerabilities of RRAM-based accelerators hinder their real application. Current research on safeguarding RRAM-based accelerators predominantly relies on a single-level protection approach. This has resulted in the restriction of its protection scope, the rigidity and lack of generality in the protection method, or has had an impact on the computational efficiency of the system. As a result, it encounters substantial challenges in attaining comprehensive optimization across multiple dimensions, such as universality, the scope of protection, and security-related overheads. In this paper, we develop specific analyses on accelerators and attacks and build graph-based representations. Based on these, we partition the RRAM-based accelerators’ security into two levels: on-chip security and off-chip security. Furthermore, we propose a two-level protection framework for RRAM-based accelerators, which is calledOmniGuard. At the on-chip security level,OmniGuardproposes a bit-grained shuffle to achieve protection while using lightweight Benes Networks to maintain the CIM capability. At the off-chip level,OmniGuardproposes an RRAM-based AES engine to introduce the AES algorithm into the accelerator with significant acceleration and minimal overhead. Evaluation results demonstrate thatOmniGuardprovides powerful and flexible protection while achieving 1.33×∼4.38× speedup and 1.26×∼2.65× power savings, with only 5% energy overhead and 3% area overhead. Ling Liang 0003, Yunfan Yang, Jinlong Lin, Meng Li 0004, Zongwei Wang 0001, Yimao Cai |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2026 | REF-CIM: A 40-nm Non-Ideality Tolerant and Energy Efficient RRAM Compute-in-Memory Macro With Configurable Precision for Edge AI
Hao Ding 0011, Yunfan Yang, Zongwei Wang 0001, Jinshan Li, Lin Bao, Ling Liang 0003, Yimao Cai |
IEEE Trans. Circuits Syst. I Regul. Pap. | 6 |
| 2025 | Towards Effective and Sparse Adversarial Attack on Spiking Neural Networks via Breaking Invisible Surrogate GradientsabstractSpiking neural networks (SNNs) have shown their competence in handling spatial-temporal event-based data with low energy consumption. Similar to conventional artificial neural networks (ANNs), SNNs are also vulnerable to gradient-based adversarial attacks, wherein gradients are calculated by spatial-temporal back-propagation (STBP) and surrogate gradients (SGs). However, the SGs may be invisible for an inference-only model as they do not influence the inference results, and current gradient-based attacks are ineffective for binary dynamic images captured by the dynamic vision sensor (DVS). While some approaches addressed the issue of invisible SGs through universal SGs, their SGs lack a correlation with the victim model, resulting in sub-optimal performance. Moreover, the imperceptibility of existing SNN-based binary attacks is still insufficient. In this paper, we introduce an innovative potential-dependent surrogate gradient (PDSG) method to establish a robust connection between the SG and the model, thereby enhancing the adaptability of adversarial attacks across various models with invisible SGs. Additionally, we propose the sparse dynamic attack (SDA) to effectively attack binary dynamic images. Utilizing a generation-reduction paradigm, SDA can fully optimize the sparsity of adversarial perturbations. Experimental results demonstrate that our PDSG and SDA outperform state-of-the-art SNN-based attacks across various models and datasets. Specifically, our PDSG achieves 100% attack success rate on ImageNet, and our SDA obtains 82% attack success rate by modifying only 0.24% of the pixels on CIFAR10DVS. The code is available at https://github.com/ryime/PDSG-SDA. Li Lun, Kunyu Feng, Qinglong Ni, Ling Liang 0003, Yuan Wang 0001, Ying Li 0056, Dunshan Yu, Xiaoxin Cui |
CVPR | 4 |
| 2025 | HybriMoE: Hybrid CPU-GPU Scheduling and Cache Management for Efficient MoE InferenceabstractThe Mixture of Experts (MoE) architecture has demonstrated significant advantages as it enables to increase the model capacity without a proportional increase in computation. However, the large MoE model size still introduces substantial memory demands, which usually requires expert offloading on resource-constrained platforms and incurs significant overhead. Hybrid CPU-GPU inference has been proposed to leverage CPU computation to reduce expert loading overhead but faces major challenges: on one hand, the expert activation patterns of MoE models are highly unstable, rendering the fixed mapping strategies in existing works inefficient; on the other hand, the hybrid CPU-GPU schedule for MoE is inherently complex due to the diverse expert sizes, structures, uneven workload distribution, etc. To address these challenges, in this paper, we propose HybriMoE, a hybrid CPU-GPU inference framework that improves resource utilization through a novel CPU-GPU scheduling and cache management system. HybriMoE introduces (i) a dynamic intra-layer scheduling strategy to balance workloads across CPU and GPU, (ii) an impact-driven inter-layer prefetching algorithm, and (iii) a score-based caching algorithm to mitigate expert activation instability. We implement HybriMoE on top of the kTransformers framework and evaluate it on three widely used MoE-based LLMs. Experimental results demonstrate that HybriMoE achieves an average speedup of $\mathbf{1. 3 3} \times$ in the prefill stage and $1.70 \times$ in the decode stage compared to state-of-the-art hybrid MoE inference framework. Our code is available at: https://github.com/PKU-SEC-Lab/HybriMoE. Shuzhang Zhong, Yanfan Sun, Ling Liang 0003, Runsheng Wang, Ru Huang 0001, Meng Li 0004 |
DAC | 3 |
| 2025 | FLASH: An Efficient Hardware Accelerator Leveraging Approximate and Sparse FFT for Homomorphic EncryptionabstractPrivate convolutional neural network (CNN) inference based on hybrid homomorphic encryption (HE) and two-party computation (2$P$C) emerges as a promising technique for sensitive user data protection. However, homomorphic convolutions (HConvs) suffer from high computation costs due to the extensive number theoretic transforms (NTTs). While customized accelerators have been proposed, they usually overlook the intrinsic error resilience and native sparsity of DNNs and hybrid HE/2$P$C protocols. In this paper, we propose FLASH, leveraging these key characteristics for highly efficient HConv. Specifically, we observe the private DNN inference is robust to computation errors and propose approximate fast Fourier transforms (FFTs) to replace NTTs and avoid the expensive modular reduction operations. We also design a flexible sparse FFT dataflow leveraging the high sparsity of weight plaintexts. With extensive experiments, we demonstrate FLASH improves the power efficiency by 90.7× for weight transforms and by 9.7× for all transforms in HConvs compared to existing works. As for the HConvs in ResNet-18 and ResNet-50, FLASH achieves about 87.3% energy consumption reduction. Ling Liang 0003, Yuan Wang 0001, Runsheng Wang, Ru Huang 0001, Meng Li 0004 |
DATE | 3 |
| 2025 | Entropy-Adaptive Diffusion Policy Optimization with Dynamic Step Alignment
Renye Yan, Jikang Cheng, Yaozhong Gan, Shikun Sun, Yunfan Yang, Ling Liang 0003, Jinlong Lin, Yeshuang Zhu, Jie Zhou 0001, Junliang Xing, Yimao Cai, Ru Huang 0001 |
ICCV | 7 |
| 2025 | SA-CIM: A 28nm 16Mb RRAM-based Sparsity-Aware Compute-In-Memory Macro for Edge AI Algorithm ProcessingabstractCompute-in-memory (CIM) for edge devices is usually constrained by on-chip resources, including on-chip memory and physical chip size, which hinders the deployment of more complex neural networks. By leveraging the sparsity of neural networks, the overall memory requirements and energy consumption can be reduced. However, Existing sparsity-aware architectures cannot achieve high energy efficiency due to off-chip sparsity control. This work proposes:1) Hybrid sparsity regulation strategy. The sparsity encoding and alignment circuit is designed and implemented, realizing on-chip sparsity detecting and encoding. 2) Sparsity-aware compute-in-memory (CIM) array based on RRAMs. The in-situ deployment of unstructured sparsity is implemented inside the CIM array, and the CIM array and sparsity are tightly coupled by sparsity read/write. This work demonstrates the design and evaluation of SA-CIM: a sparsity-aware CIM macro with 16Mb RRAM with fine-grained sparsity detecting and encoding capacity, achieving energy efficiency of 22.7TOP/W@8b/8b. Hao Ding 0011, Zongwei Wang 0001, Jinshan Li, Shigeng Zhao, Heting Gao, Junbo Ao, Ling Liang 0003, Yimao Cai, Ru Huang 0001 |
ISCAS | 7 |
| 2025 | Ironman: Accelerating Oblivious Transfer Extension for Privacy-Preserving AI with Near-Memory Processing
Chenqi Lin, Kang Yang 0002, Tianshi Xu, Ling Liang 0003, Runsheng Wang, Mingyu Gao 0001, Meng Li 0004 |
MICRO | 4 |
| 2025 | Matrix: Multi-Cipher Structures Dataflow for Parallel and Pipelined TFHE AcceleratorabstractFully homomorphic encryption over torus (TFHE) enables the execution of arbitrary functions on encrypted data through programmable bootstrapping (PBS). However, performing all operations on ciphertext during PBS results in high computational and memory requirements, limiting the deployment of PBS in real-world scenarios. Previous TFHE accelerator designs have attempted to improve performance by employing specific dataflow and functional units, but these techniques may require large off-chip bandwidth or on-chip storage when scaling up computation capacity. Additionally, the design of specialized functional units may limit the utilization of computation units when facing dynamic secure parameter settings. To address these challenges and further improve PBS throughput in TFHE, we propose Matrix , an ASIC-based architecture that balances off-chip bandwidth and on-chip storage according to the execution flow of PBS. In Matrix , we utilize a unified special-prime-based processing element (PE) that achieves high utilization with minimal resource overhead. Furthermore, we propose a hybrid PBS dataflow that can efficiently reduce computation complexity and memory requirements. Compared to state-of-the-art TFHE accelerators, Matrix achieves 1.43 × -5.66 × throughput improvement for PBS. For ZAMA Deep-NN benchmark, we achieve 525.60× and 68.06× speedup compared to CPU and GPU, respectively. 1 Ling Liang 0003, Fahong Zhang 0004, Zhirui Li, Xin Fan 0009, Dimin Niu, Meng Li 0004, Zhiyong Li 0016, Zongwei Wang 0001, Hongzhong Zheng, Yimao Cai, Yuan Xie 0001 |
ACM Trans. Archit. Code Optim. | 1 |
| 2024 | TensorTEE: Unifying Heterogeneous TEE Granularity for Efficient Secure Collaborative Tensor ComputingabstractHeterogeneous collaborative computing with NPU and CPU has received widespread attention due to its substantial performance benefits. To ensure data confidentiality and integrity during computing, Trusted Execution Environments (TEE) is considered a promising solution because of its comparatively lower overhead. However, existing heterogeneous TEE designs are inefficient for collaborative computing due to fine and different memory granularities between CPU and NPU. 1) The cacheline granularity of CPU TEE intensifies memory pressure due to its extra memory access, and 2) the cacheline granularity MAC of NPU escalates the pressure on the limited memory storage. 3) Data transfer across heterogeneous enclaves relies on the transit of non-secure regions, resulting in cumbersome re-encryption and scheduling. Husheng Han, Xinyao Zheng, Yuanbo Wen 0001, Yifan Hao 0001, Erhu Feng, Ling Liang 0003, Jianan Mu, Xiaqing Li, Tianyun Ma, Pengwei Jin, Xinkai Song, Zidong Du, Qi Guo 0001, Xing Hu 0001 |
ASPLOS (4) | 6 |
| 2024 | A High-Throughput Private Inference Engine Based on 3D Stacked MemoryabstractFully Homomorphic Encryption (FHE) enables unlimited computation depth, allowing privacy-enhanced neural network inference tasks directly on the ciphertext. However, existing FHE architectures suffer from the memory access bottleneck. This work proposes a High-throughput FHE engine for private inference (PI) based on 3D stacked memory (H3). H3 adopts the software-hardware co-design that dynamically adjusts the polynomial decomposition during the PI process to minimize the computation and storage overhead at a fine granularity. With 3D hybrid bonding, H3 integrates a logic die with a multi-layer embedded DRAM, routing data efficiently to the processing unit array through an efficient broadcast mechanism. H3 consumes 192mm2 when implemented using a 28nm logic process. It achieves 1.36 million LeNet-5 or 920 ResNet-20 PI per minute, surpassing existing 7nm accelerators by 52%. This demonstrates that 3D memory is a promising technology to promote the performance of FHE. Ling Liang 0003, Zhirui Li, Fahong Zhang 0004, Yanheng Lu |
DAC | 2 |
| 2024 | AdapMoE: Adaptive Sensitivity-based Expert Gating and Management for Efficient MoE InferenceabstractMixture-of-Experts (MoE) models are designed to enhance the efficiency of large language models (LLMs) without proportionally increasing the computational demands. However, their deployment on edge devices still faces significant challenges due to high on-demand loading overheads from managing sparsely activated experts. This paper introduces AdapMoE, an algorithm-system co-design framework for efficient MoE inference. AdapMoE features adaptive expert gating and management to reduce the on-demand loading overheads. We observe the heterogeneity of experts loading across layers and tokens, based on which we propose a sensitivity-based strategy to adjust the number of activated experts dynamically. Meanwhile, we also integrate advanced prefetching and cache management techniques to further reduce the loading latency. Through comprehensive evaluations on various platforms, we demonstrate AdapMoE consistently outperforms existing techniques, reducing the average number of activated experts by 25% and achieving a 1.35× speedup without accuracy degradation. Code is available at: https://github.com/PKU-SEC-Lab/AdapMoE. Shuzhang Zhong, Ling Liang 0003, Yuan Wang 0001, Runsheng Wang, Ru Huang 0001, Meng Li 0004 |
ICCAD | 2 |
| 2024 | Autoencoder Reconstruction Model for Long-Horizon ExplorationabstractConventional reinforcement learning (RL) algorithms often necessitate millions of environment interactions to ascertain an efficacious policy. In stark contrast, humans, leveraging their curiosity mechanisms, can develop proficient policies with minimal effort. Drawing inspiration from this observation, we introduce the Autoencoder Reconstruction Model(ARM), a curiosity-driven RL model that significantly reduces interactions while enhancing policy effectiveness. ARM employs an autoencoder module, utilizing a deep neural network to learn feature representations from the environment. ARM utilizes its Curiosity Measurement Module to motivate RL agents for effective exploration, particularly in environments with sparse rewards. ARM also introduces an innovative mechanism to balance the exploration-exploitation dilemma. Theoretical analyses reveal that the reward shaping introduced by the ARM aligns with the potential-based reward shaping paradigm, thereby preserving the optimality of reinforcement learning. We will release the source code and trained models to facilitate further studies in this research direction. Renye Yan, Yaozhong Gan, Yunfan Yang, Zhaoke Yu, Zongxi Liu, Ling Liang 0003, Yimao Cai |
IJCNN | 8 |
| 2023 | SPG: Structure-Private Graph Database via SqueezePIRabstractMany relational data in our daily life are represented as graphs, making graph application an important workload. Because of the large scale of graph datasets, moving graph data to the cloud becomes a popular option. To keep the confidential and private graph secure from an untrusted cloud server, many cryptographic techniques are leveraged to hide the content of the data. However, protecting only the data content is not enough for a graph database. Because the structural information of the graph can be revealed through the database accessing track. In this work, we study the graph neural network (GNN), an important graph workload to mine information from a graph database. We find that the server is able to infer which node is processing during the edge retrieving phase and also learn its neighbor indices during GNN's aggregation phase. This leads to the leakage of the information of graph structure data. In this work, we present SPG, a structure-private graph database with SqueezePIR. Our SPG is built on top of Private Information Retrieval (PIR), which securely hides which nodes/neighbors are accessed. In addition, we propose SqueezePIR, a compression technique to overcome the computation overhead of PIR. Based on our evaluation, our SqueezePIR achieves 11.85× speedup on average with less than 2% accuracy loss when compared to the state-of-the-art FastPIR protocol. Ling Liang 0003, Jilan Lin, Zheng Qu 0002, Ishtiyaque Ahmad, Fengbin Tu, Trinabh Gupta, Yufei Ding 0001, Yuan Xie 0001 |
Proc. VLDB Endow. | 1 |
| 2023 | SDP: Co-Designing Algorithm, Dataflow, and Architecture for In-SRAM Sparse NN AccelerationabstractProcessing-in-memory (PIM) is a promising architecture for neural network (NN) acceleration. Most previous PIMs are based on analog computing, so their accuracy and memory cell array utilization are limited by analog deviation and ADC overhead. Digital PIM is an emerging type of PIM architecture that integrates digital logic in memory cells, which can make full utilization of the cell array without accuracy loss. However, digital PIM’s rigid crossbar architecture and full array activation raise new challenges in sparse NN acceleration. Conventional unstructured or structured sparsity cannot perform well on both the weight and input side of digital PIM. We take the opportunities from digital PIM’s bit-serial processing and in-memory customization, to tackle the above challenges by the co-designing sparse algorithm, multiplication dataflow, and PIM architecture. At the algorithm level, we propose double-broadcast hybrid-grained pruning to exploit weight sparsity with better accuracy and efficiency balance. At the dataflow level, we propose a bit-serial Booth in-SRAM multiplication dataflow for stable acceleration from the input side. At the architecture level, we design a sparse digital PIM (SDP) accelerator with customized SRAM-PIM macros to support the proposed techniques. SDP achieves$3.59\times $,$8.15\times $,$3.11\times $area efficiency, and$6.95\times $,$29.44\times $,$39.40\times $energy savings, over state-of-the-art sparse NN architectures SIGMA, SRE, and Bit Prudent. Fengbin Tu, Yiqi Wang 0005, Ling Liang 0003, Yufei Ding 0001, Leibo Liu, Shaojun Wei, Shouyi Yin, Yuan Xie 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2023 | Comprehensive SNN Compression Using ADMM Optimization and Activity RegularizationabstractAs well known, the huge memory and compute costs of both artificial neural networks (ANNs) and spiking neural networks (SNNs) greatly hinder their deployment on edge devices with high efficiency. Model compression has been proposed as a promising technique to improve the running efficiency via parameter and operation reduction, whereas this technique is mainly practiced in ANNs rather than SNNs. It is interesting to answer how much an SNN model can be compressed without compromising its functionality, where two challenges should be addressed: 1) the accuracy of SNNs is usually sensitive to model compression, which requires an accurate compression methodology and 2) the computation of SNNs is event-driven rather than static, which produces an extra compression dimension on dynamic spikes. To this end, we realize a comprehensive SNN compression through three steps. First, we formulate the connection pruning and weight quantization as a constrained optimization problem. Second, we combine spatiotemporal backpropagation (STBP) and alternating direction method of multipliers (ADMMs) to solve the problem with minimum accuracy loss. Third, we further propose activity regularization to reduce the spike events for fewer active operations. These methods can be applied in either a single way for moderate compression or a joint way for aggressive compression. We define several quantitative metrics to evaluate the compression performance for SNNs. Our methodology is validated in pattern recognition tasks over MNIST, N-MNIST, CIFAR10, and CIFAR100 datasets, where extensive comparisons, analyses, and insights are provided. To the best of our knowledge, this is the first work that studies SNN compression in a comprehensive manner by exploiting all compressible components and achieves better results. Lei Deng 0003, Yujie Wu 0002, Yifan Hu 0013, Ling Liang 0003, Guoqi Li 0002, Xing Hu 0001, Yufei Ding 0001, Peng Li 0001, Yuan Xie 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2023 | Exploring Adversarial Attack in Spiking Neural Networks With Spike-Compatible GradientabstractSpiking neural network (SNN) is broadly deployed in neuromorphic devices to emulate brain function. In this context, SNN security becomes important while lacking in-depth investigation. To this end, we target the adversarial attack against SNNs and identify several challenges distinct from the artificial neural network (ANN) attack: 1) current adversarial attack is mainly based on gradient information that presents in a spatiotemporal pattern in SNNs, hard to obtain with conventional backpropagation algorithms; 2) the continuous gradient of the input is incompatible with the binary spiking input during gradient accumulation, hindering the generation of spike-based adversarial examples; and 3) the input gradient can be all-zeros (i.e., vanishing) sometimes due to the zero-dominant derivative of the firing function. Recently, backpropagation through time (BPTT)-inspired learning algorithms are widely introduced into SNNs to improve the performance, which brings the possibility to attack the models accurately given spatiotemporal gradient maps. We propose two approaches to address the above challenges of gradient-input incompatibility and gradient vanishing. Specifically, we design a gradient-to-spike (G2S) converter to convert continuous gradients to ternary ones compatible with spike inputs. Then, we design a restricted spike flipper (RSF) to construct ternary gradients that can randomly flip the spike inputs with a controllable turnover rate, when meeting all-zero gradients. Putting these methods together, we build an adversarial attack methodology for SNNs. Moreover, we analyze the influence of the training loss function and the firing threshold of the penultimate layer on the attack effectiveness. Extensive experiments are conducted to validate our solution. Besides the quantitative analysis of the influence factors, we also compare SNNs and ANNs against adversarial attacks under different attack methods. This work can help reveal what happens in SNN attacks and might stimulate more research on the security of SNN models and neuromorphic devices. Ling Liang 0003, Xing Hu 0001, Lei Deng 0003, Yujie Wu 0002, Guoqi Li 0002, Yufei Ding 0001, Peng Li 0001, Yuan Xie 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2022 | Accelerating Spatiotemporal Supervised Training of Large-Scale Spiking Neural Networks on GPUabstractSpiking neural networks (SNNs) have great potential to achieve brain-like intelligence, however, it suffers low accuracy of conventional synaptic plasticity rules and low training efficiency on GPUs. Recently, the emerging backpropagation through time (BPTT) inspired learning algorithms bring new opportunities to boost the accuracy of SNNs, while training on GPUs still remains inefficient due to the complex spatiotemporal dynamics and huge memory consumption, which restricts the model exploration for SNNs and prevents the advance of neuromorphic computing. In this work, we build a framework to solve the inefficiency of BPTT-based SNN training on modern GPUs. To reduce the memory consumption, we optimize the dataflow by saving CONV/FC results only in the forward pass and recomputing other intermediate results in the backward pass. Then, we customize kernel functions to accelerate the neural dynamics for all training stages. Finally, we provide a Pytorch interface to make our framework easy-to-deploy in real systems. Compared to vanilla Pytorch implementation, our framework can achieve up to 2.13 x end-to-end speedup and consume only 0.41 x peak memory on the CIFAR10 dataset. Moreover, for the distributed training on the large ImageNet dataset, we can achieve up to 1.81 x end-to-end speedup and consume only 0.38 x peak memory. Ling Liang 0003, Zhaodong Chen 0001, Lei Deng 0003, Fengbin Tu, Guoqi Li 0002, Yuan Xie 0001 |
DATE | 1 |
| 2022 | INSPIRE: in-storage private information retrieval via protocol and architecture co-designabstractPrivate Information Retrieval (PIR) plays a vital role in secure, database-centric applications. However, existing PIR protocols explore a massive working space containing hundreds of GiBs of query and database data. As a consequence, PIR performance is severely bounded by storage communication, making it far from practical for real-world deployment. Jilan Lin, Ling Liang 0003, Zheng Qu 0002, Ishtiyaque Ahmad, Liu Liu 0017, Fengbin Tu, Trinabh Gupta, Yufei Ding 0001, Yuan Xie 0001 |
ISCA | 2 |
| 2022 | Toward Robust Spiking Neural Network Against Adversarial PerturbationabstractAs spiking neural networks (SNNs) are deployed increasingly in real-world efficiency critical applications, the security concerns in SNNs attract more attention.Currently, researchers have already demonstrated an SNN can be attacked with adversarial examples. How to build a robust SNN becomes an urgent issue.Recently, many studies apply certified training in artificial neural networks (ANNs), which can improve the robustness of an NN model promisely. However, existing certifications cannot transfer to SNNs directly because of the distinct neuron behavior and input formats for SNNs. In this work, we first design S-IBP and S-CROWN that tackle the non-linear functions in SNNs' neuron modeling. Then, we formalize the boundaries for both digital and spike inputs. Finally, we demonstrate the efficiency of our proposed robust training method in different datasets and model architectures. Based on our experiment, we can achieve a maximum $37.7\%$ attack error reduction with $3.7\%$ original accuracy loss. To the best of our knowledge, this is the first analysis on robust training of SNNs. Ling Liang 0003, Kaidi Xu, Xing Hu 0001, Lei Deng 0003, Yuan Xie 0001 |
NeurIPS | 1 |
| 2022 | A Systematic View of Model Leakage Risks in Deep Neural Network SystemsabstractAs deep neural networks (DNNs) continue to find applications in ever more domains, the exact nature of the neural network architecture becomes an increasingly sensitive subject, due to either intellectual property protection or risks of adversarial attacks. While prior work has explored aspects of the risk associated with model leakage, exactly which parts of the model are most sensitive and how one infers the full architecture of the DNN when nothing is known about the structure a priori are problems that have been left unexplored. In this paper we address this gap, first by presenting a schema for reasoning about model leakage holistically, and then by proposing and quantitatively evaluating DeepSniffer, a novel learning-based model extraction framework that uses no prior knowledge of the victim model. DeepSniffer is robust to architectural and system noises introduced by the complex memory hierarchy and diverse run-time system optimizations. Taking GPU platforms as a showcase, DeepSniffer performs model extraction by learning both the architecture-level execution features of kernels and the inter-layer temporal association information introduced by the common practice of DNN design. We demonstrate that DeepSniffer works experimentally in the context of an off-the-shelf Nvidia GPU platform running a variety of DNN models and that the extracted models significantly improve attempts at crafting adversarial inputs. The DeepSniffer project has been released inhttps://github.com/xinghu7788/DeepSniffer. Xing Hu 0001, Ling Liang 0003, Xiaobing Chen, Lei Deng 0003, Yu Ji 0002, Yufei Ding 0001, Zidong Du, Qi Guo 0001, Timothy Sherwood, Yuan Xie 0001 |
IEEE Trans. Computers | 2 |
| 2022 | Rubik: A Hierarchical Architecture for Efficient Graph Neural Network TrainingabstractThe graph convolutional network (GCN) emerges as a promising direction to learn the inductive representation in graph data commonly used in widespread applications, such as E-commerce, social networks, and knowledge graphs. However, learning from graphs is nontrivial because of its mixed computation model involving both graph analytics and neural network computing. To this end, we decompose the GCN learning into two hierarchical paradigms: 1) graph-level and 2) node-level computing. Such a hierarchical paradigm facilitates the software and hardware accelerations for GCN learning. We propose a lightweight graph reordering methodology, incorporated with a GCN accelerator architecture that equips a customized cache design to fully utilize the graph-level data reuse. We also propose a mapping methodology aware of data reuse and task-level parallelism to handle various graphs inputs effectively. The results show that Rubik accelerator design improves energy efficiency by$26.3\times $–$1375.2\times $than GPU platforms across different datasets and GCN models. Xiaobing Chen, Xinfeng Xie, Xing Hu 0001, Abanti Basak, Ling Liang 0003, Mingyu Yan, Lei Deng 0003, Yufei Ding 0001, Zidong Du, Yuan Xie 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 6 |
| 2022 | H2Learn: High-Efficiency Learning Accelerator for High-Accuracy Spiking Neural NetworksabstractAlthough spiking neural networks (SNNs) take benefits from the bioplausible neural modeling, the low accuracy under the common local synaptic plasticity learning rules limits their application in many practical tasks. Recently, an emerging SNN supervised learning algorithm inspired by backpropagation through time (BPTT) from the domain of artificial neural networks (ANNs) has successfully boosted the accuracy of SNNs, and helped improve the practicability of SNNs. However, current general-purpose processors suffer from low efficiency when performing BPTT for SNNs due to the ANN-tailored optimization. On the other hand, current neuromorphic chips cannot support BPTT because they mainly adopt local synaptic plasticity rules for simplified implementation. In this work, we propose H2Learn, a novel architecture that can achieve high efficiency for BPTT-based SNN learning, which ensures high accuracy of SNNs. At the beginning, we characterized the behaviors of BPTT-based SNN learning. Benefited from the binary spike-based computation in the forward pass and weight update, we first design look-up table (LUT)-based processing elements in the forward engine and weight update engine to make accumulations implicit and to fuse the computations of multiple input points. Second, benefited from the rich sparsity in the backward pass, we design a dual-sparsity-aware backward engine, which exploits both input and output sparsity. Finally, we apply a pipeline optimization between different engines to build an end-to-end solution for the BPTT-based SNN learning. Compared with the modern NVIDIA V100 GPU, H2Learn achieves$7.38\times $area saving,$5.74-10.20\times $speedup, and$5.25-7.12\times $energy saving on several benchmark datasets. Ling Liang 0003, Zheng Qu 0002, Zhaodong Chen 0001, Fengbin Tu, Yujie Wu 0002, Lei Deng 0003, Guoqi Li 0002, Peng Li 0001, Yuan Xie 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2022 | Hardware-Enabled Efficient Data Processing With Tensor-Train DecompositionabstractIn recent years, tensor computation has become a promising tool for solving big data analysis, machine learning, medical image, and EDA problems. To ease the memory and computation intensity of tensor processing, decomposition techniques, especially tensor-train decomposition (TTD), are widely adopted to compress the extremely high-dimensional tensor data. Despite TTD’s potential to break the curse of dimensionality, researchers have not yet leveraged its full computational potential, mainly because of two reasons: 1) executing TTD itself is time- and energy-consuming due to the singular value decomposition (SVD) operation inside each of TTD’s iteration and 2) additional software/hardware optimizations are often required to process the obtained TT-format data in certain applications such as deep learning inference. In this article, we address these challenges with two approaches. First, we propose an algorithm-hardware co-design with customized architecture, namely, TTD Engine to accelerate TTD. We use MRI image compression as a demo application to illustrate the efficacy of the proposed accelerator. Second, we present a case study demonstrating the benefit of TT-format data processing and the efficacy of using TTD Engine. In the case study, we use the TT approach to realize convolution operation, which is difficult and nontrivial for TT-format data. Experimental results show that, TTD Engine achieves, on average,$14.9 \times $–$36.9 \times $speedup over CPU implementations and$4.1\times $–$9.9\times $speedup compared to the GPU baseline. The energy efficiency is also improved by at least$14.4\times $and$5.4\times $over CPU and GPU, respectively. Moreover, our hardware-enabled TT-format data processing further leads to more efficient implementations of complicated operations and applications. Zheng Qu 0002, Lei Deng 0003, Bangyan Wang, Hengnu Chen, Jilan Lin, Ling Liang 0003, Guoqi Li 0002, Zheng Zhang 0005, Yuan Xie 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 6 |
| 2021 | SEALing Neural Network Models in Encrypted Deep Learning AcceleratorsabstractDeep learning (DL) accelerators suffer from a new security problem, i.e., being vulnerable to physical access based attacks. An adversary can easily obtain the entire neural network (NN) model by physically snooping the memory bus that connects the accelerator chip with DRAM memory. Therefore, memory encryption becomes important for DL accelerators to improve their security. Nevertheless, we observe that traditional memory encryption techniques that have been efficiently used in CPU systems cause significant performance degradation when directly used in DL accelerators, due to the big bandwidth gap between the memory bus and the encryption engine. To address this problem, our paper proposes SEAL, a Secure and Efficient Accelerator scheme for deep Learning to enhance the performance of encrypted DL accelerators by improving the data access bandwidth. Specifically, SEAL leverages a criticality-aware smart encryption scheme that identifies partial data having no impact on the security of NN models and allows them to bypass the encryption engine, thus reducing the amount of data to be encrypted without affecting security. Extensive experimental results demonstrate that, compared with existing memory encryption techniques, SEAL achieves 1.34 – 1.4× overall performance improvement. Pengfei Zuo, Yu Hua 0001, Ling Liang 0003, Xinfeng Xie, Xing Hu 0001, Yuan Xie 0001 |
DAC | 3 |
| 2021 | SpaceA: Sparse Matrix Vector Multiplication on Processing-in-Memory AcceleratorabstractSparse matrix-vector multiplication (SpMV) is an important primitive across a wide range of application domains such as scientific computing and graph analytics. Due to its intrinsic memory-bound characteristics, the performance of SpMV on throughput-oriented architectures such as GPU is bounded by the limited bandwidth between processors and memory. Processing-in-memory (PIM) architectures, made feasible by advances in 3D stacking, provide new opportunities to utilize ultra-high bandwidth by integrating compute-logic into memory.In this paper, we develop an SpMV accelerator, named as SpaceA, based on PIM architectures. SpaceA integrates compute logic near memory banks to exploit bank-level bandwidth. SpaceA contains both hardware and data-mapping design features to alleviate irregular memory access patterns which hinder full utilization of high memory bandwidth. In terms of hardware design features, SpaceA consists of two unique features: (1) it utilizes the capability of outstanding memory requests to hide the memory access latency to data located in non-local memory banks; (2) it integrates Content Addressable Memory (CAM) at the bank level to exploit data reuse of the input vectors. In addition, we develop a mapping scheme that partitions the sparse matrix into different memory banks, to maximize the data locality of the input vector and to achieve workload balance among processing elements (PEs) near each bank. Overall, SpaceA together with the proposed mapping method achieves 13.54x speedup and 87.49% energy saving on average over the GPU baseline on SpMV computation. In addition to SpMV primitives, we conduct a case study on graph analytics to demonstrate the benefits of SpaceA for applications built on SpMV. Compared to Tesseract and GraphP, state-of-the-art graph accelerators, SpaceA obtains better performance due to its higher effective bandwidth provided by near-bank integration. Xinfeng Xie, Zheng Liang 0003, Peng Gu 0008, Abanti Basak, Lei Deng 0003, Ling Liang 0003, Xing Hu 0001, Yuan Xie 0001 |
HPCA | 6 |
| 2021 | Brain-Inspired Computing: Adventure from Beyond CMOS Technologies to Beyond von Neumann Architectures ICCAD Special Session PaperabstractThe goal of this special session paper is to introduce and discuss different breakthrough technologies as well as novel architectures and how they together may reshape the future of Artificial Intelligent. Our aim is to provide a comprehensive overview on the latest advances in brain-inspired computing and how the latter can be realized when emerging technologies, using beyond-CMOS devices, are coupled with novel computing paradigms that go beyond von Neumann architectures. Different emerging technologies like Ferroelectric Field-Effect Transistor (FeFET), Phase Change Memory (PCM), and Resistive RAM (ReRAM) are discussed, demonstrating their promising capability in building neuromorphic computing architectures that are inspired by nature. In addition, this special session paper discusses various novel concepts such as Logic-in-Memory (LIM), Processing-in-Memory (PIM), and Spiking Neural Networks (SNNs) towards exploring the far-reaching consequences of beyond von Neumann computing on accelerating deep learning. Finally, the latest trends in brain-inspired computing are summarized into algorithm, technology, and application-driven innovations towards comparing different PIM architectures. Hussam Amrouch, Jian-Jia Chen, Kaushik Roy 0001, Yuan Xie 0001, Indranil Chakraborty, Wenqin Huangfu, Ling Liang 0003, Fengbin Tu, Cheng Wang 0036, Mikail Yayla |
ICCAD | 7 |
| 2021 | ScaleCert: Scalable Certified Defense against Adversarial Patches with Sparse Superficial LayersabstractAdversarial patch attacks that craft the pixels in a confined region of the input images show their powerful attack effectiveness in physical environments even with noises or deformations. Existing certified defenses towards adversarial patch attacks work well on small images like MNIST and CIFAR-10 datasets, but achieve very poor certified accuracy on higher-resolution images like ImageNet. It is urgent to design both robust and effective defenses against such a practical and harmful attack in industry-level larger images. In this work, we propose the certified defense methodology that achieves high provable robustness for high-resolution images and largely improves the practicality for real adoption of the certified defense. The basic insight of our work is that the adversarial patch intends to leverage localized superficial important neurons (SIN) to manipulate the prediction results. Hence, we leverage the SIN-based DNN compression techniques to significantly improve the certified accuracy, by reducing the adversarial region searching overhead and filtering the prediction noises. Our experimental results show that the certified accuracy is increased from 36.3% (the state-of-the-art certified detection) to 60.4%on the ImageNet dataset, largely pushing the certified defenses for practical use. Husheng Han, Kaidi Xu, Xing Hu 0001, Xiaobing Chen, Ling Liang 0003, Zidong Du, Qi Guo 0001, Yanzhi Wang 0001, Yunji Chen |
NeurIPS | 5 |
| 2021 | Tensor train decomposition for solving large-scale linear equations
Hengnu Chen, Lei Deng 0003, Zheng Qu 0002, Ling Liang 0003, Tianyi Yan, Yuan Xie 0001, Guoqi Li 0002 |
Neurocomputing | 4 |
| 2021 | Practical Attacks on Deep Neural Networks by Memory TrojaningabstractDeep neural network (DNN) accelerators are widely deployed in computer vision, speech recognition, and machine translation applications, in which attacks on DNNs have become a growing concern. This article focuses on exploring the implications of hardware Trojan attacks on DNNs. Trojans are one of the most challenging threat models in hardware security where adversaries insert malicious modifications to the original integrated circuits (ICs), leading to malfunction once being triggered. Such attacks can be conducted by adversaries because modern ICs commonly include third-party intellectual property (IP) blocks. Previous studies design hardware Trojans to attack DNNs with the assumption that adversaries have full knowledge or manipulation of the DNN systems' victim model and toolchain in addition to the hardware platforms, yet such a threat model is strict, limiting their practical adoption. In this article, we propose a memory Trojan methodology that implants the malicious logics merely into the memory controllers of DNN systems without the necessity of toolchain manipulation or accessing to the victim model and thus is feasible for practical uses. Specifically, we locate the input image data among the massive volume of memory traffics based on memory access patterns and propose a Trojan trigger mechanism based on detecting the geometric feature in input images. Extensive experiments show that the proposed trigger mechanism is effective even in the presence of environmental noises and preprocessing operations. Furthermore, we design and implement the payload and verify that the proposed Trojan technique can effectively conduct both untargeted and targeted attacks on DNNs. Xing Hu 0001, Yang Zhao 0013, Lei Deng 0003, Ling Liang 0003, Pengfei Zuo, Jing Ye 0001, Yingyan (Celine) Lin, Yuan Xie 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2021 | Effective and Efficient Batch Normalization Using a Few Uncorrelated Data for Statistics EstimationabstractDeep neural networks (DNNs) thrive in recent years, wherein batch normalization (BN) plays an indispensable role. However, it has been observed that BN is costly due to the huge reduction and elementwise operations that are hard to be executed in parallel, which heavily reduces the training speed. To address this issue, in this article, we propose a methodology to alleviate the BN's cost by using only a few sampled or generated data for mean and variance estimation at each iteration. The key challenge to reach this goal is how to achieve a satisfactory balance between normalization effectiveness and execution efficiency. We identify that the effectiveness expects less data correlation in sampling while the efficiency expects more regular execution patterns. To this end, we design two categories of approach: sampling or creating a few uncorrelated data for statistics' estimation with certain strategy constraints. The former includes "batch sampling (BS)" that randomly selects a few samples from each batch and "feature sampling (FS)" that randomly selects a small patch from each feature map of all samples, and the latter is "virtual data set normalization (VDN)" that generates a few synthetic random samples to directly create uncorrelated data for statistics' estimation. Accordingly, multiway strategies are designed to reduce the data correlation for accurate estimation and optimize the execution pattern for running acceleration in the meantime. The proposed methods are comprehensively evaluated on various DNN models, where the loss of model accuracy and the convergence rate are negligible. Without the support of any specialized libraries, 1.98× BN layer acceleration and 23.2% overall training speedup can be practically achieved on modern GPUs. Furthermore, our methods demonstrate powerful performance when solving the well-known "micro-BN" problem in the case of a tiny batch size. This article provides a promising solution for the efficient training of high-performance DNNs. Zhaodong Chen 0001, Lei Deng 0003, Guoqi Li 0002, Xing Hu 0001, Ling Liang 0003, Yufei Ding 0001, Yuan Xie 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 6 |
| 2020 | DeepSniffer: A DNN Model Extraction Framework Based on Learning Architectural HintsabstractAs deep neural networks (DNNs) continue their reach into a wide range of application domains, the neural network architecture of DNN models becomes an increasingly sensitive subject, due to either intellectual property protection or risks of adversarial attacks. Previous studies explore to leverage architecture-level events disposed in hardware platforms to extract the model architecture information. They pose the following limitations: requiring a priori knowledge of victim models, lacking in robustness and generality, or obtaining incomplete information of the victim model architecture. Xing Hu 0001, Ling Liang 0003, Shuangchen Li, Lei Deng 0003, Pengfei Zuo, Yu Ji 0002, Xinfeng Xie, Yufei Ding 0001, Chang Liu 0021, Timothy Sherwood, Yuan Xie 0001 |
ASPLOS | 2 |
| 2020 | HyGCN: A GCN Accelerator with Hybrid ArchitectureabstractInspired by the great success of neural networks, graph convolutional neural networks (GCNs) are proposed to analyze graph data. GCNs mainly include two phases with distinct execution patterns. The Aggregation phase, behaves as graph processing, showing a dynamic and irregular execution pattern. The Combination phase, acts more like the neural networks, presenting a static and regular execution pattern. The hybrid execution patterns of GCNs require a design that alleviates irregularity and exploits regularity. Moreover, to achieve higher performance and energy efficiency, the design needs to leverage the high intra-vertex parallelism in Aggregation phase, the highly reusable inter-vertex data in Combination phase, and the opportunity to fuse phase-by-phase execution introduced by the new features of GCNs. However, existing architectures fail to address these demands. In this work, we first characterize the hybrid execution patterns of GCNs on Intel Xeon CPU. Guided by the characterization, we design a GCN accelerator, HyGCN, using a hybrid architecture to efficiently perform GCNs. Specifically, first, we build a new programming model to exploit the fine-grained parallelism for our hardware design. Second, we propose a hardware design with two efficient processing engines to alleviate the irregularity of Aggregation phase and leverage the regularity of Combination phase. Besides, these engines can exploit various parallelism and reuse highly reusable data efficiently. Third, we optimize the overall system via inter-engine pipeline for inter-phase fusion and priority-based off-chip memory access coordination to improve off-chip bandwidth utilization. Compared to the state-of-the-art software framework running on Intel Xeon CPU and NVIDIA V100 GPU, our work achieves on average 1509× speedup with 2500× energy reduction and average 6.5× speedup with 10× energy reduction, respectively. Mingyu Yan, Lei Deng 0003, Xing Hu 0001, Ling Liang 0003, Yujing Feng, Xiaochun Ye, Zhimin Zhang 0004, Dongrui Fan, Yuan Xie 0001 |
HPCA | 4 |
| 2020 | Rethinking the performance comparison between SNNS and ANNS
Lei Deng 0003, Yujie Wu 0002, Xing Hu 0001, Ling Liang 0003, Yufei Ding 0001, Guoqi Li 0002, Guang-She Zhao, Peng Li 0001, Yuan Xie 0001 |
Neural Networks | 4 |
| 2020 | SemiMap: A Semi-Folded Convolution Mapping for Speed-Overhead Balance on CrossbarsabstractCrossbar architecture has been widely used in neural network (NN) accelerators, involving conventional and emerging devices. It performs well on the fully connected layer through efficient vector-matrix multiplication. Whereas, the advantages degrade on the convolutional layer with huge data reuse, since the execution speed and resource overhead are imbalanced when using existing fully unfolded or fully folded mapping strategy. To address this issue, we propose a novel semi-folded mapping (SemiMap) framework for implementing the convolution on crossbars. It simultaneously folds the physical resources along the row dimension of feature maps (FMs) and unfolds them along the column dimension. The former reduces the resource overhead, and the latter maintains the parallelism. An FM slicing scheme is further proposed to enable the processing of large-size image. Via our mapping framework, a row-by-row streaming pipeline for intraimage dataflow and periodical pipeline for interimage dataflow are easy to be obtained. To validate the idea, we build a many-crossbar architecture with several designs to guarantee the overall functionality and performance. Based on the measurement data of a fabricated chip, a mapping compiler and a cycle-accurate simulator are developed for the hardware simulation of large-scale networks. We evaluate the proposed SemiMap on various convolutional NNs across different network scale. ${>} 35 {\times }$ resource saving and several hundred times cycle reduction are demonstrated compared to the existing fully unfolded and fully folded strategies, respectively. This paper jumps out of the current extreme mapping schemes, and provides a balanced solution on how to efficiently deploy the computational graphs with data reuse on many-crossbar architecture. Lei Deng 0003, Yuan Xie 0001, Ling Liang 0003, Guanrui Wang, Liang Chang 0002, Xing Hu 0001, Liu Liu 0017, Jing Pei, Guoqi Li 0002 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2018 | TETRIS: TilE-matching the TRemendous Irregular SparsityabstractCompressing neural networks by pruning weights with small magnitudes can significantly reduce the computation and storage cost. Although pruning makes the model smaller, it is difficult to get practical speedup in modern computing platforms such as CPU and GPU due to the irregularity. Structural pruning has attract a lot of research interest to make sparsity hardware-friendly. Increasing the sparsity granularity can lead to better hardware utilization, but it will compromise the sparsity for maintaining accuracy. In this work, we propose a novel method, TETRIS, to achieve both better hardware utilization and higher sparsity. Just like a tile-matching game, we cluster the irregularly distributed weights with small value into structured groups by reordering the input/output dimension and structurally prune them. Results show that it can achieve comparable sparsity with the irregular element-wise pruning and demonstrate negligible accuracy loss. The experiments also shows ideal speedup, which is proportional to the sparsity, on GPU platforms. Our proposed method provides a new solution toward algorithm and architecture co-optimization for accuracy-efficiency trade-off. Yu Ji 0002, Ling Liang 0003, Lei Deng 0003, Youyang Zhang, Youhui Zhang, Yuan Xie 0001 |
NeurIPS | 2 |