Lei Wang 0011

dblp:w/LeiWang11 · DBLP profile ↗
← Back
52ranked-venue papers
3as first author
31since 2021 · last 2026
0000-0003-3249-7333ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 38 · 2 first-author · 23 since 2021Artificial intelligence and machine learning · 7 · 4 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-author · 1 since 2021Software engineering, systems software and programming languages · 2 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021Security and privacy · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1
YearPublicationVenuePosition
2026 EstCoder: A RTL Code Generator based on Static Functional Estimation
abstract
Optimizing register transfer level (RTL) code is of vital importance in hardware design. Large language models (LLMs) provide new methods for the automatic generation and optimization of RTL code. However, existing methods for generating RTL code often focus on model fine-tuning and the use of various expansion techniques to enhance the RTL code generation capabilities, lacking attention to the functional correctness. To address this issue, we propose EstCoder, an LLM-powered collaborative agent framework for RTL code generation based on static functional score estimation. EstCoder operates a three-stage paradigm: Generation, Estimation and Correction. During the stages, the functional estimation agent statically evaluates the generated code based on score and assessment results, and decides whether to output the code directly, return it for regeneration, or forward it to the code correction agent. This famework can be applied to various LLMs that designed for RTL code generation, further enhancing the correctness of the generated code. By providing quantitative scores and human-readable requirements comparisons, it improves the transparency of AI-assisted RTL code generation. Experiments show that EstCoder significantly improves the correctness of RTL code generation by generic LLM by 3.2%-9.0%, demonstrating the practical value of our system.
Renzhi Chen, Zhigang Fang 0002, Bowei Wang, Libo Huang 0002, Lei Wang 0011
DATE7
2026 A hardware-efficient FPGA-based YOLOv5 accelerator with operator fusion and unified dataflow scheduling
Libo Huang 0002, Run Yan, Lei Wang 0011, Jianzhuang Lu
Future Gener. Comput. Syst.4
2026 NPEva: An NoC-Based Neuromorphic Processors Performance Evaluation Framework for Benchmarking Spiking Neural Networks
abstract
Spiking Neural Network (SNN) applications place diverse demands on neuromorphic processors’ computational and communication capabilities. Specifically, the sparse parallel computation of SNNs requires storage systems capable of extensive parallel data access and processing. Additionally, the time-dependent nature of SNN computations demands a communication framework to manage spike transmission and synchronization. Addressing these requires continuous early-stage evaluation of potential designs, which can be time-intensive. To this end, this paper introduces a performance evaluation model for NoC-based neuromorphic processors, named NPEva, which includes a Computation Model (CpMo) and a Communication Model (CoMo). This model facilitates rapid, high-dimensional exploration of design spaces across a wide range of microarchitectural parameters. CpMo quickly assesses computation-related latency, power consumption, and area by extracting relevant parameters from hierarchical storage organizations and neuron computation models. CoMo employs communication upper-bound latency analysis to evaluate packet latencies to reduce redundant synchronization time. Comparisons with TrueNorth’s publicly released data reveal that the model’s evaluation results are within an 8% error margin. The CoMo offers enhanced precision in estimating latency upper-bounds compared to other models like worst-contention delay (WCD) and worst-case traversal time (WCTT), especially when varying the number of VCs. With VC settings of 2, 4, 8, and 16, CoMo reduced latency by 0.55× to 2.53× compared to WCTT and by 0.83× to 25.57× compared to WCD. Additionally, using the NPEva model for exploring the design space of Liquid State Machines (LSM) networks has shown that optimized hardware can improve cost efficiency by 2.94× with only a 1% increase in latency.
Ziyang Kang, Lei Wang 0011, De Ma, Gang Pan 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3
2026 A2SPM: A Memory Access Acceleration for Neuromorphic Processing with the SPM
abstract
Owing to their low power consumption and high energy efficiency, neuromorphic processors have found extensive applications across various intelligent computing domains, including artificial intelligence (AI) and low-power edge computing. Homogeneous neuromorphic processors demonstrate a unique integration of general-purpose computing versatility and neuromorphic computing acceleration capabilities. These processors effectively address the intensive memory access demands inherent in neuromorphic computing through the implementation of tightly coupled local memory, specifically Scratchpad Memory (SPM). This article presents the design of SPM-based memory access acceleration mechanisms, which have been specifically developed and optimized to accommodate the data characteristics and computational processes of typical Spiking Neural Networks (SNNs). The proposed scheme significantly enhances computational efficiency by reducing both memory access operations and computational instruction processing. Comparative performance evaluations reveal that the proposed design achieves a 28.99% speedup in Liquid State Machine operations and a 14.95% speedup in Spiking Convolutional Neural Network computations when benchmarked against a baseline homogeneous processor equipped with SPM.
Xun Xiao, Lei Wang 0011
ACM Trans. Embed. Comput. Syst.7
2025 VToT: Automatic Verilog Generation via LLMs with Tree of Thoughts Prompting
abstract
The automatic generation of Verilog code using Large Language Models (LLMs) presents a compelling solution to enhance the efficiency of hardware design flow. However, the state-of-the-art performance of LLMs in Verilog generation remains limited compared to programming languages such as Python. Previous research, Chain of Thought (CoT), has demonstrated that incorporating intermediate reasoning steps can significantly improve the performance of LLMs in code generation. In this paper, we propose the Verilog Tree of Thoughts (VToT) method. This structured prompting technique addresses the abstraction gap between Verilog and CoT by embedding hierarchical design constraints within the prompt. Experimental results on the VerilogEval and RTLLM benchmarks demonstrate that VToT prompting enhances both the syntactic and functional correctness of the generated code. Specifically, according to the RTLLM benchmark, VToT achieved a correctness rate of 75.9% at pass@5, representing an improvement of 10.4%. Furthermore, in the VerilogEval benchmark, VToT achieved state-of-the-art performance with a correctness rate of 52.4% at pass@1 (an increase of 8.9%) and 65.4% at pass@5 (an increase of 9.6%).
Renzhi Chen, Zhigang Fang 0002, Bowei Wang, Wenqiang Bai, Qilin Cao, Lei Wang 0011
DATE9
2025 LintLLM: An Open-Source Verilog Linting Framework Based on Large Language Models
Zhigang Fang 0002, Renzhi Chen, Yang Guo 0003, Huadong Dai, Lei Wang 0011
ACM Great Lakes Symposium on VLSI6
2025 SICA: A Multicore Neuromorphic Processor Featuring Sparse Integration and Communication-Aware Optimization
abstract
Neuromorphic computing has emerged as a promising paradigm due to its event-driven operation and energy efficiency, driving extensive research in neuromorphic processor development. When implementing spiking neural networks (SNNs) on such processors, two critical aspects must be addressed: neuron computation and spike communication. For neuron computation, previous work primarily relies on parallel accumulation via adder trees but fails to leverage the inherent sparsity in SNNs. For spike communication, conventional mesh topologies suffer from long-distance communication inefficiencies, while suboptimal mapping strategies further exacerbate latency issues. To address these challenges, we propose a low-overhead fast sparse detection mechanism that effectively exploits spike sparsity and optimizes the processor's workflow, thereby achieving efficient synaptic integration with minimal overhead. For spike communication, we employ an on-chip broadcast mechanism combined with a hybrid torus-mesh topology to significantly reduce communication latency, while systematically evaluating the impact of three distinct mapping strategies-random, sequential, and communication-aware mapping-on overall performance. Experimental results demonstrate significant improvements, with our solution delivering speedups of$1.26 \times$and$1.24 \times$compared to LSMCore on the N-MNIST and MNIST datasets, respectively. Furthermore, the communication-aware mapping strategy achieves a 24.87% reduction in communication latency, while the torus topology contributes an additional$\mathbf{1 4. 4 8} \boldsymbol{\%}$latency reduction.
Junbo Tie, Xun Xiao, Yuanfeng Luo, Yang Guo 0003, Lei Wang 0011
HPCC13
2025 RTLBench: A Multi-Dimensional Benchmark Suite for Evaluating LLM-Generated RTL Code
abstract
The rapid advancement of large language models (LLMs) has enabled automated Register Transfer Level (RTL) code generation, accelerating chip design workflows. However, existing benchmarks focus mainly on syntax and functionality, overlooking critical engineering aspects such as lint compliance, readability, and coding style. To address this gap, we propose RTLBench, a benchmark suite of 160 copyright-free RTL cases sourced from textbooks and open-source projects. RTLBench features a multi-dimensional evaluation framework covering syntax, functionality, lint compliance, readability, and style consistency. To assess subjective code quality metrics, it also incorporates an LLM-as-a-judge mechanism. We evaluated 24 state-of-the-art LLMs using RTLBench, finding that while several models perform well in syntax and functionality, most fall short on engineering quality. To address this, we propose Log2BetterRTL, a log-driven feedback system that transforms EDA tool diagnostics into iterative improvement prompts. It improves syntax correctness by up to 18.13 %, boosts functional correctness by 14.38 %, reduces lint violations by up to 229, and raises clarity scores by 0.51. These results demonstrate RTLBench's effectiveness in evaluating and enhancing LLMgenerated RTL, bridging the gap between generative AI and industrial-grade hardware design. The suite and scripts are available at: https://fangzhigang32.github.io/RTLBench.
Zhigang Fang 0002, Renzhi Chen, Yang Guo 0003, Huadong Dai, Lei Wang 0011
ICCD5
2025 An NoC-Based Latency Upper-Bound Model for Reducing Timestep Length in SNN Communication
abstract
Spiking Neural Networks (SNNs) with real-time demands are deployed on Network-on-Chip (NoC)-based neuromorphic processors for specific tasks. While meeting real-time constraints for spiking data streams is crucial, overly long timesteps (e.g., 1ms) result in idle neuron cores and routers, reducing efficiency. This paper proposes a communication performance model using a recursive calculation method to assess the worst-case upper-bound latency. Integrated with the NoC router microarchitecture, the model analyzes the behavior of spiking streams during communication. It effectively balances timestep length to ensure real-time communication while minimizing idle time. Empirical results show that our model outperforms worst-contention delay (WCD) and worst-case traversal time (WCTT) models, achieving latency reductions of 0.55× to 2.53× compared to WCTT and 0.83× to 25.57× compared to WCD across various spiking datasets and virtual channels. Encouragingly, with only a 1% accuracy reduction, latency was reduced by 18%.
Ziyang Kang, Lei Wang 0011, De Ma, Gang Pan 0001
ISCAS3
2025 Hardware/Software Co-design for spike communication optimization: Leveraging neuron-level communication patterns
Yuanfeng Luo, Weixia Xu 0001, Lei Wang 0011
J. Syst. Archit.7
2025 EBF: An Event-Based Bilateral Filter for Effective Neuromorphic Vision Sensor Denoising
Shasha Guo 0001, Chenyang Shi, Lei Wang 0011, Yuliang Lu
IEEE Trans. Circuits Syst. Video Technol.3
2024 A security JPEG image system accelerated by NEON technology based on FT-2000/4
Junbo Tie, Lei Wang 0011
CCF Trans. High Perform. Comput.5
2024 A robust defense for spiking neural networks against adversarial examples via input filtering
Shasha Guo 0001, Lei Wang 0011, Yuliang Lu
J. Syst. Archit.2
2024 Hierarchical Mapping of Large-Scale Spiking Convolutional Neural Networks Onto Resource-Constrained Neuromorphic Processor
abstract
Neuromorphic processors have been designed as non-von Neumann systems for energy-efficient spiking neural network (SNN) execution. Spiking convolutional neural networks (SCNNs), combining the advantage of convolutional neural network (CNN) and SNN, have been widely applied to vision tasks. However, as the scale of SCNNs increases, executing large-scale SCNNs on resource-constrained neuromorphic processor faces many challenges, including massive synapse pruning caused by resource competition, execution performance degradation, etc. Addressing these problems, we propose an efficient approach to map large-scale SCNNs onto resource-constrined neuromorphic processor. The approach consists of three steps: splitting, partitioning, and mapping. We explore three acyclic splitting strategies to divide large-scale SCNNs into subnetworks without cyclic dependency. Axon sharing is the guiding principle to partition subnetworks into multiple clusters. To obtain an optimal cluster-to-core mapping scheme, we use Non-dominated Sorting Genetic Algorithm to collaboratively optimize two metrics. We evaluate our approach with eight realistic SCNN applications. The results show that compared with existing state-of-the-art methods, our approach significantly reduces the synapse pruning and accuracy loss, and increases the execution performance.
Xun Xiao, Yao Wang 0002, Junbo Tie, Lei Wang 0011, Weixia Xu 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.8
2024 LSM-Based Hotspot Prediction and Hotspot-Aware Routing in NoC-Based Neuromorphic Processor
abstract
The traffic patterns of spiking neural networks (SNNs) exhibit high variability and stochastic, leading to the emergence of elevated traffic hotspots on the network-on-chip (NoC)-based neuromorphic processors. Predicting the occurrence of hotspots remains one of the most challenging issues in NoC design. This article presents the first attempt toward traffic hotspot prediction by utilizing liquid state machine (HP-LSM). The predictor extracts essential information reflecting the current state of the NoC to predict potential routing hotspots in the subsequent time step. Furthermore, we designed the hardware architecture for HP-LSM, which incorporates leaky-integrate-and-fire (LIF) neurons with configurable biological parameters. Meanwhile, we introduce a novel hotspot-aware path-based multicast (HaPM) routing algorithm that utilizes advanced knowledge acquired from HP-LSM to guide packet routing throughout the network, aiming to improve the performance of NoC. Results indicate that the HP-LSM can forecast hotspot formation with an accuracy up to 89.36% and 90.19% for two spiking-based datasets, respectively. The hardware experiment results demonstrate a 92.03% reduction in the average execution time of zero skipping compared with nonzero skipping. Moreover, the HP-LSM exhibits a reduction of up to 79.30% in the number of neurons compared with other related SNN predictor models. The experiments reveal a reduction of 73.67% and 53.42% in the average length of the multicast path when compared with dual-path (DP) or multipath (MP) multicast routing. The HaPM demonstrates improved performance in terms of average latency and throughput compared with DP, MP, and path-based multicast (PbM) multicast routing.
Ziyang Kang, Xun Xiao, Lei Wang 0011, De Ma, Gang Pan 0001
IEEE Trans. Very Large Scale Integr. Syst.5
2023 Dynamic Obstacle Avoidance for Unmanned Aerial Vehicle Using Dynamic Vision Sensor
Junbo Tie, Jingyue Zhao, Zhong Wan, Guangda Zhang, Lei Wang 0011
ICANN (10)12
2023 M-LSM: An Improved Multi-Liquid State Machine for Event-Based Vision Recognition
Lei Wang 0011, Shasha Guo 0001, Lianhua Qu, Shuo Tian, Weixia Xu 0001
J. Comput. Sci. Technol.1
2023 Back to Homogeneous Computing: A Tightly-Coupled Neuromorphic Processor With Neuromorphic ISA
abstract
In recent years, neuromorphic processors are widely used in many scenarios, showing extreme energy efficiency over traditional architectures. However, almost all existing neuromorphic hardware are following the heterogeneous computing methodology without Instruction Set Architecture (ISA), leading to inflexibility in programming. In this paper, we first propose a RISC-V Neuromorphic Extension (RVNE) to enable fine-grained and flexible homogeneous programming for neuromorphic algorithms while utilizing SNN sparsity from different levels of granularity and computing flows. Based on RVNE, we next implement a neuromorphic micro-architecture that is tightly coupled to the CPU pipeline to accelerate neuromorphic computing. To demonstrate the proposed homogeneous neuromorphic architecture, we implement a prototype processor called NeuroRVcore based on RISC-V ISA and an open-source RISC-V core. The evaluation results show that RVNE achieves a 2.8 × −4.3 × reduction in code density compared with the general-purpose ISAs. Compared with the state-of-the-art neuromorphic processor, the proposed homogeneous computing reduces energy consumption by 3.4%−22.5% while enabling fine-grained and flexible homogeneous programming.
Lei Wang 0011, Yao Wang 0002, Junbo Tie, Feng Wang 0050, LingHui Peng, Xun Xiao, Gan Zhou, Xuhu Yu, Xia Zhao 0004, Yuhua Tang, Weixia Xu 0001
IEEE Trans. Parallel Distributed Syst.2
2022 Unicorn: a multicore neuromorphic processor with flexible fan-in and unconstrained fan-out for neurons
abstract
Neuromorphic processor is popular due to its high energy efficiency for spatio-temporal applications. However, when running the spiking neural network (SNN) topologies with the ever-growing scale, existing neuromorphic architectures face challenges due to their restrictions on neuron fan-in and fan-out. This paper proposes Unicorn, a multicore neuromorphic processor with a spike train sliding multicasting mechanism (STSM) and neuron merging mechanism (NMM) to support unconstrained fan-out and flexible fan-in of neurons. Unicorn supports 36K neurons and 45M synapses and thus supports a variety of neuromorphic applications. The peak performance and energy efficiency of Unicorn reach 36TSOPS and 424GSOPS/W respectively. Experimental results show that Unicorn can achieve 2×-5.5× energy reduction over the state-of-the-art neuromorphic processor when running an SNN with a relatively large fan-out and fan-in.
Lei Wang 0011, Yao Wang 0002, LingHui Peng, Xun Xiao, Weixia Xu 0001
DAC2
2022 Dynamic Vision Sensor Based Gesture Recognition Using Liquid State Machine
Xun Xiao, Lei Wang 0011, Lianhua Qu, Shasha Guo 0001, Yao Wang 0002, Ziyang Kang
ICANN (3)2
2022 A Spatio-Temporal Event Data Augmentation Method for Dynamic Vision Sensor
Xun Xiao, Ziyang Kang, Shasha Guo 0001, Lei Wang 0011
ICONIP (6)5
2022 NeuProMa: A Toolchain for Mapping Large-Scale Spiking Convolutional Neural Networks onto Neuromorphic Processor
Jihua Chen, Lei Wang 0011
NPC3
2022 A Hardware Security Isolation Architecture for Intelligent Accelerator
abstract
AI systems face potential hardware security threats. Existing AI systems generally use the heterogeneous architecture of CPU + Intelligent Accelerator, with PCIe bus for communication between them. Security mechanisms are implemented on CPUs based on the hardware security isolation architecture. But the conventional hardware security isolation architecture does not include the intelligent accelerator on the PCIe bus. Therefore, from the perspective of hardware security, data offloaded to the intelligent accelerator face great security risks. In order to effectively integrate intelligent accelerator into the CPU’s security mechanism, a novel hardware security isolation architecture is presented in this paper. The PCIe protocol is extended to be security-aware by adding security information packaging and unpacking logic in the PCIe controller. The hardware resources on the intelligent accelerator are isolated in fine-grained. The resources classified into the secure world can only be controlled and used by the software of CPU’s trusted execution environment. Based on the above hardware security isolation architecture, a security isolation spiking convolutional neural network accelerator is designed and implemented in this paper. The experimental results demonstrate that the proposed security isolation architecture has no overhead on the bandwidth and latency of the PCIe controller. The architecture does not affect the performance of the entire hardware computing process from CPU data offloading, intelligent accelerator computing, to data returning to CPU. With low hardware overhead, this security isolation architecture achieves effective isolation and protection of input data, model, and output data. And this architecture can effectively integrate hardware resources of intelligent accelerator into CPU’s security isolation mechanism.
Lei Wang 0011
TrustCom2
2022 Hardware-aware liquid state machine generation for 2D/3D Network-on-Chip platforms
Ziyang Kang, Lei Wang 0011, Lianhua Qu, Weixia Xu 0001
J. Syst. Archit.3
2022 LSMCore: A 69k-Synapse/mm2 Single-Core Digital Neuromorphic Processor for Liquid State Machine
abstract
Neuromorphic processors have gained momentum recently due to their high energy efficiency in artificial intelligence applications compared to DNN accelerators. Most neuromorphic processors are executing SNNs (Spiking Neural Networks). Liquid State Machine (LSM), as the spiking version of reservoir computing, shows advantages and great potential in image classification, speech recognition, language translation, etc.. Comparing with other SNN models, LSM has the characteristics of easy to train and low resource utilization, which is suitable for low-power and resource-constrained edge computing scenarios. In this paper, we propose a novel design of a neuromorphic processor, LSMCore, aiming at LSM acceleration. LSMCore supports both training and inference of LSM. It consists of 256 input neurons, 1024 liquid neurons, and 1.31M synapses. Besides, multiple optimization techniques, including weight quantization for reducing storage, zero-skipping for decreasing dynamic sparsity, and mini-batch training are adopted in this processor. The experimental results show that the frequency of LSMCore achieves 400 MHz, the power is 4.9W and the area is 18.49 mm2with a 40nm library. Comparing with the baseline, LSMCore achieves up to$80.7\times $($49.6\times $),$91.3\times $($56.3\times $), and$83.1\times $($56.8\times $) speedup on MNIST, N-MNIST, and Free Spoken Digital Dataset (FSDD) respectively for training (inference), while the accuracy of LSMCore on these three datasets are 96.8%, 97.6%, and 90% respectively.
Lei Wang 0011, Shasha Guo 0001, Lianhua Qu, Ziyang Kang, Weixia Xu 0001
IEEE Trans. Circuits Syst. I Regul. Pap.1
2021 A Novel Ring-based Small-World NoC for Neuromorphic Processor
abstract
Neuromorphic computing has shown promise in metrics such as power consumption and parallelism over existing computer systems, which essentially promote the development of neuromorphic processors in recent years. In order to properly place the increasing number of neuron cores and support the inter-core communication, Network-on-Chip (NoC) is widely used in the design of neuromorphic processors. Mesh has historically been used for multi-core NoCs, however, in neuromorphic chips, computation cores are relatively small, mesh-based SNN with high resource occupation limited the peak performance and energy efficiency. Moreover, one of the most significant findings in the neuroscience is that human brain exhibits small-world effect which is originated from the social network, inspired by that, we proposed a composite architecture for neuromorphic processor called the ring-based small-world NoC. Correspondingly, we proposed a routing algorithm to by generating a specific-application routing table in the pre-processing stage. We evaluated the performance such as delay, energy and resource utilization comprehensively based on three spike-based datasets (FSDD, NMNIST and N-TIDIGITS). The experimental results show that the average packet latency and the resource utilization of the proposed network is reduced by up to 18% and 35%, compared to a regular mesh network.
Yuchen Qiu, LingHui Peng, Ziyang Kang, Lei Wang 0011
ASAP7
2021 A Hardware Aware Liquid State Machine Generation Framework
abstract
The liquid state machine (LSM) is a kind of spiking neural network (SNN) that usually is mapped to an NoC-based neuromorphic processor to perform tasks such as classification. The creation of these LSM models does not consider the structure of Network on Chip (NoC) which resulting in heavy communication pressure on the NoC. In this paper, we propose a hardware aware LSM network generation framework. By keeping the communication between neurons within cores as much as possible, this framework could reduce the communication overheads between cores effectively. The experimental results show that the LSM model produced by our framework could achieve state-of-art accuracy and is hardware-friendly. Compared with the mapping method, the synapses in our LSM is reduced by 94.14%, the total packets in NoC is reduced by 78.3%, the maximum transmission latency is reduced by 97.5%, the average transmission latency is reduced by 54%, the throughput increased 2.8x.
Ziyang Kang, Lei Wang 0011, Lianhua Qu
ISCAS3
2021 Fine-Grained Video Deblurring with Event Camera
Limeng Zhang, Chenyang Zhu 0002, Shasha Guo 0001, Jihua Chen, Lei Wang 0011
MMM (1)6
2021 A neural architecture search based framework for liquid state machine design
Shuo Tian, Lianhua Qu, Lei Wang 0011, Weixia Xu 0001
Neurocomputing3
2021 HashHeat: A hashing-based spatiotemporal filter for dynamic vision sensor
Shasha Guo 0001, Ziyang Kang, Lei Wang 0011, Limeng Zhang, Weixia Xu 0001
Integr.3
2021 A multi-objective LSM/NoC architecture co-design framework
Shuo Tian, Ziyang Kang, Lianhua Qu, Lei Wang 0011, Weixia Xu 0001
J. Syst. Archit.6
2020 HashHeat: An O(C) Complexity Hashing-based Filter for Dynamic Vision Sensor
abstract
Neuromorphic event-based dynamic vision sensors (DVS) have much faster sampling rates and a higher dynamic range than frame-based imagers. However, they are sensitive to background activity (BA) events which are unwanted. We propose HashHeat, a hashing-based BA filter with O(C) complexity. It is the first spatiotemporal filter that doesn't scale with the DVS output size N and doesn't store the 32-bits timestamps. HashHeat consumes 100x less memory and increases the signal to noise ratio by 15x compared to previous designs.
Shasha Guo 0001, Ziyang Kang, Lei Wang 0011, Weixia Xu 0001
ASP-DAC3
2020 Application-specific network-on-chip design space exploration framework for neuromorphic processor
abstract
Neuromorphic processors can support the design of various Spiking Neural Networks (SNN) to deal with different tasks, such as recognition and tracking. Neuromorphic processors use Network-on-Chip (NoC) to support communication between neurons in SNN. The different SNN has different communication traffic patterns. It will pose the different challenges of the NoC designing. A reasonable NoC architecture can improve the overall performance such as lower latency of the processor. Hence, it is critical to implement the exploration of NoC architecture design for neuromorphic processors.
Ziyang Kang, Lei Wang 0011, Lianhua Qu, Weixia Xu 0001
CF3
2020 SNEAP: A Fast and Efficient Toolchain for Mapping Large-Scale Spiking Neural Network onto NoC-based Neuromorphic Platform
abstract
Spiking neural network (SNN), as the third generation of artificial neural networks, has been widely adopted in vision and audio tasks. Nowadays, many neuromorphic platforms support SNN simulation and adopt Network-on-Chips (NoC) architecture for multi-cores interconnection. However, a large volume and run-time communication on the interconnection has a significant effect on performance of the platform. In this paper, we propose a toolchain called SNEAP (Spiking NEural network mAPping toolchain) for mapping SNNs to neuromorphic platforms with multi-cores, which aims to reduce the energy and latency brought by spike communication on the interconnection.
Shasha Guo 0001, Limeng Zhang, Ziyang Kang, Lei Wang 0011, Weixia Xu 0001
ACM Great Lakes Symposium on VLSI7
2020 Real-Time Gesture Classification System Based on Dynamic Vision Sensor
Limeng Zhang, Shasha Guo 0001, Lianhua Qu, Lei Wang 0011
ICONIP (1)6
2020 FTR-NAS: Fault-Tolerant Recurrent Neural Architecture Search
Shuo Tian, Lei Wang 0011
ICONIP (5)6
2020 Recurrent Neural Architecture Search based on Randomness-Enhanced Tabu Algorithm
abstract
Deep neural networks have achieved highly competitive performance in multiple tasks in recent years. However, discovering state-of-the-art neural network architectures requires substantial effort from human experts. To speed up the process, neural architecture search (NAS) has been proposed to search promising architectures automatically. Nevertheless, the search process of NAS is computing-expensive and time-consuming, which even costs thousands of GPU days. In this paper, to solve the bottleneck, we apply the randomness-enhanced tabu algorithm as a controller to sample candidate architectures, which balances the global exploration and local exploitation for the architectural solutions. In addition, more aggressive weight-sharing strategy is introduced into our method, which significantly reduces the overhead of evaluating sampled architectures. Our approach discovers the recurrent neural architecture within 0.78 GPU hour, which is 15.3x more efficient than ENAS [1] in terms of search time, and the architecture we discovered achieves the test perplexity of 56.1 on Penn Tree Bank (PTB) dataset, which is lower than ENAS by 2.2. In addition, we further demonstrate the usefulness of the learned architecture by transferring it to wiki-text-2 (WT2) dataset well. Moreover, the extended experiments on the WT2 dataset also show promising results.
Shuo Tian, Shasha Guo 0001, Lei Wang 0011
IJCNN6
2020 CompressedCache: Enabling Storage Compression on Neuromorphic Processor for Liquid State Machine
Lianhua Qu, Ziyang Kang, Lei Wang 0011, Weixia Xu 0001
NPC6
2020 SIES: A Novel Implementation of Spiking Convolutional Neural Network Inference Engine on Field-Programmable Gate Array
Shuquan Wang, Lei Wang 0011, Yu Deng 0001, Shasha Guo 0001, Ziyang Kang, Yu-Feng Guo, Weixia Xu 0001
J. Comput. Sci. Technol.2
2020 A Machine Learning Framework with Feature Selection for Floorplan Acceleration in IC Physical Design
Shuzheng Zhang, Zhen-Yu Zhao, Chaochao Feng, Lei Wang 0011
J. Comput. Sci. Technol.4
2020 ASIE: An Asynchronous SNN Inference Engine for AER Events Processing
abstract
Neuromorphic computing based on spiking neural network (SNN) shows good energy-efficiency. However, it is inefficient for SNN to perform the convolution based on frame. It may contain a lot of redundant information in the frame. The output of Dynamic Vision Sensors (DVS) is a stream event based on Address Event Representation (AER). The asynchronous nature of AER events makes the event-based convolution reflect the characteristics of SNN low energy consumption. This article presents an SNN hardware inference engine based on an asynchronous Processing Element (PE) array with AER events as input. The engine uses a convolution algorithm based on AER events. This design also uses distributed storage in the PE array to store the state of neurons to reduce the cost of memory access. The experimental results show that the design can achieve a recognition accuracy of 98.0% for the MNIST AER dataset. The design can perform the reference process more efficiently in the case where the accuracy of the loss is negligible. During the filling and draining processes of the systolic array, the number of active PE units in our PE array is reduced and, thus, the average power consumption per PE unit is drastically decreased.
Ziyang Kang, Lei Wang 0011, Shasha Guo 0001, Yu Deng 0001, Weixia Xu 0001
ACM J. Emerg. Technol. Comput. Syst.2
2020 CSMO-DSE: Fast and Precise Application-driven DSE Guided by Criticality and Sensitivity Analysis
abstract
Determining the optimal microarchitecture configuration of a processor at the early stages of design is undeniably a challenge. Due to many parameters at the microarchitecture level, finding the proper combination of these parameters to arrive at a balanced design is difficult. Application-specific Design Space Exploration (DSE) is even more difficult, since the property of application needs to be considered during the DSE process. Improving the speed and accuracy of the DSE process remains a particular challenge in microprocessor design. In this article, we propose a novel processor DSE methodology based on criticality and sensitivity analysis, named Criticality and Sensitivity-based Multi-Objective DSE (CSMO-DSE). In our methodology, a dependence-graph is derived from the profile generated by running a program on an instrumented cycle-accurate microprocessor simulator. Then, the criticality of the processor’s performance events is obtained through critical path analysis. The sensitivity of microarchitecture parameters to various performance events is also analyzed. Then, this information is used to optimize performance, power/area, and energy efficiency of the design. Experiments with SPEC 2006 show that CSMO-DSE methodology is 4.73× faster than the baseline DSE methodology and that the quality of result (QoR) is better than the baseline methodology for all the benchmark programs.
Lei Wang 0011, Yu Deng 0001, Yongwen Wang
ACM J. Emerg. Technol. Comput. Syst.1
2019 A Systolic SNN Inference Accelerator and its Co-optimized Software Framework
abstract
Although Deep Neural Network (DNN) architectures have made some breakthroughs in computer vision tasks, they are not close to biological brain neurons. Spiking Neural Network (SNN) is highly expected to bridge the gap between artificial computing systems and bio-systems. And it also shows great potential in low power computing. This paper presents a low power hardware accelerator for SNN inference using systolic array, and a corresponding software framework for optimization. First, we give the hardware design which adopts systolic array inspired by explorations of SNN. Then we ensure correct data mapping for systolic array for the sake of computational correctness. Next, we use compression methods for decreasing both the runtime and memory footprint. Finally, we make the systolic array size-configurable to adapt to different input, so as to reduce computational overhead. We implement the accelerator on Xilinx FPGA V7 690T. The experimental results show that SNN inference on our scheme suffers little loss on accuracy (less than 0.1%) on MNIST and Fashion-MNIST, and the runtime of the time-consuming layers decreases. The total power of our scheme is 0.745 W at 100 MHz.
Shasha Guo 0001, Lei Wang 0011, Shuquan Wang, Yu Deng 0001, Zhige Xie, Qiang Dou
ACM Great Lakes Symposium on VLSI2
2019 PRTSM: Hardware Data Arrangement Mechanisms for Convolutional Layer Computation on the Systolic Array
Shuquan Wang, Lei Wang 0011, Shuo Tian, Shasha Guo 0001, Ziyang Kang, Shuzheng Zhang, Weixia Xu 0001
NPC2
2018 Live Demonstration: Image Segmentation on the FPGA-based Pre-calculating Ising Memory
abstract
We demonstrate image segmentation processing by using a pre-calculating Ising memory implemented on a FPGA. Results show that the FPGA-based pre-calculating Ising memory can segment a prepared image in under 100μs.
Jian Zhang 0022, Shuming Chen, Lei Wang 0011, Linghui Lv
ISCAS4
2018 Pre-Calculating Ising Memory: Low Cost Method to Enhance Traditional Memory with Ising Ability
abstract
Combinatorial optimization always contains many state search operations, which greatly reduce the efficiency of Von Neumann architecture. The Ising chip, expressing the behavior of magnetic spin systems with CMOS circuit, can efficiently support such operations. On the Ising chip, the state search can be carried out for all the spins in parallel. As the Ising chip is mainly SRAM based architecture, we propose Ising memory that enhancing the traditional memory with Ising ability, which can be easily integrated into Von Neumann architecture for both traditional data storage and efficiently solving combinatorial optimization problems. However, due to the non-memory logic for state search operations, directly integrating Ising ability into traditional memory would introduce additional 2× area overhead. To solve this problem, we propose pre-calculating structure to reduce the complexity of the state search circuit. Our proposal helps to reduce the non-memory area overhead to about 0.9× of the traditional memory. Moreover, we have physically designed an Ising memory and tested it with image segmentation problems. Our Ising memory can accelerate the segmenting processing by 26000× with only 0.0001× energy consumption. The experiment result shows that our Ising memory is a low cost method to enhance traditional memory with Ising ability for both data storage and solving combinatorial optimization problems.
Jian Zhang 0022, Shuming Chen, Lei Wang 0011, Linghui Lv
ISCAS4
2018 Systolic Array Based Accelerator and Algorithm Mapping for Deep Learning Algorithms
Lei Wang 0011, Yu Deng 0001, Qiang Dou
NPC2
2017 FixCaffe: Training CNN with Low Precision Arithmetic Operations by Fixed Point Caffe
Shasha Guo 0001, Lei Wang 0011, Baozi Chen, Qiang Dou, Yuxing Tang, Zhisheng Li
APPT2
2016 The Macro-DSE for HPC Processing Unit: The Physical Constraints Perspective
Yuxing Tang, Lei Wang 0011, Yu Deng 0001, Xiaoqiang Ni, Qiang Dou
GPC2
2016 Shielding STT-RAM Based Register Files on GPUs against Read Disturbance
abstract
To address the high energy consumption issue of SRAM on GPUs, emerging Spin-Transfer Torque (STT-RAM) memory technology has been intensively studied to build GPU register files for better energy-efficiency, thanks to its benefits of low leakage power, high density, and good scalability. However, STT-RAM suffers from the read disturbance issue, which stems from the fact that the voltage difference between read current and write current becomes smaller as technology scales. The read disturbance leads to high error rates for read operations, which cannot be effectively protected by the SEC-DED ECC on large-capacity register files of GPUs. Prior schemes (e.g., read-restore) to mitigate the read disturbance usually incur either non-trivial performance loss or excessive energy overhead, thus not applicable for the GPU register file design that aims to achieve both high performance and energy-efficiency. To combat the read disturbance, we propose a novel software-hardware co-designed solution (i.e., Red-Shield ), which consists of three optimizations to overcome the limitations of the existing solutions. First, we identify dead reads at compiling stage and augment instructions to avoid unnecessary restores. Second, we employ a small read buffer to accommodate register reads with high-access locality to further reduce restores. Third, we propose an adaptive restore mechanism to selectively pick the suitable restore scheme, according to the busy status of corresponding register banks. Experimental results show that our proposed design can effectively mitigate the performance loss and energy overhead caused by restore operations while still maintaining the reliability of reads.
Xuhao Chen 0001, Nong Xiao 0001, Lei Wang 0011, Fang Liu 0002, Wei Chen 0009, Zhiguang Chen 0001
ACM J. Emerg. Technol. Comput. Syst.4
2009 The Design of Asynchronous Microprocessor Based on Optimized NCL_X Design-Flow
abstract
NCL X circuit is a very efficient way to implement the QDI circuit, which can get all the advantages of the asynchronous circuit, especially the average performance. But the NCL_X circuits suffer from its huge area overhead. To solve this problem, a method for optimizing the complete detection network in the NCL_X circuit has been introduced in this paper. Using this method can dramatically reduced the area of the NCL_X circuit, according to the experimental result, the area of the NCL_X circuit may be reduced more than 60%. We also use this optimized method to implement an asynchronous microprocessor pipeline (APC). Compared to the synchronous implementation, this NCL_X implementation can achieve higher performance because the NCL_X circuit can get the average performance.
Gang Jin, Lei Wang 0011, Zhiying Wang 0003
NAS2
2007 An Optimal Design Method for De-synchronous Circuit Based on Control Graph
Gang Jin, Lei Wang 0011, Zhiying Wang 0003, Kui Dai
APPT2