LingHui Peng

dblp:300/5276 · DBLP profile ↗
← Back
4ranked-venue papers
0as first author
4since 2021 · last 2023
0009-0008-1109-8492ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 4 · 4 since 2021
YearPublicationVenuePosition
2023 Back to Homogeneous Computing: A Tightly-Coupled Neuromorphic Processor With Neuromorphic ISA
abstract
In recent years, neuromorphic processors are widely used in many scenarios, showing extreme energy efficiency over traditional architectures. However, almost all existing neuromorphic hardware are following the heterogeneous computing methodology without Instruction Set Architecture (ISA), leading to inflexibility in programming. In this paper, we first propose a RISC-V Neuromorphic Extension (RVNE) to enable fine-grained and flexible homogeneous programming for neuromorphic algorithms while utilizing SNN sparsity from different levels of granularity and computing flows. Based on RVNE, we next implement a neuromorphic micro-architecture that is tightly coupled to the CPU pipeline to accelerate neuromorphic computing. To demonstrate the proposed homogeneous neuromorphic architecture, we implement a prototype processor called NeuroRVcore based on RISC-V ISA and an open-source RISC-V core. The evaluation results show that RVNE achieves a 2.8 × −4.3 × reduction in code density compared with the general-purpose ISAs. Compared with the state-of-the-art neuromorphic processor, the proposed homogeneous computing reduces energy consumption by 3.4%−22.5% while enabling fine-grained and flexible homogeneous programming.
Lei Wang 0011, Yao Wang 0002, Junbo Tie, Feng Wang 0050, LingHui Peng, Xun Xiao, Gan Zhou, Xuhu Yu, Xia Zhao 0004, Yuhua Tang, Weixia Xu 0001
IEEE Trans. Parallel Distributed Syst.8
2022 Unicorn: a multicore neuromorphic processor with flexible fan-in and unconstrained fan-out for neurons
abstract
Neuromorphic processor is popular due to its high energy efficiency for spatio-temporal applications. However, when running the spiking neural network (SNN) topologies with the ever-growing scale, existing neuromorphic architectures face challenges due to their restrictions on neuron fan-in and fan-out. This paper proposes Unicorn, a multicore neuromorphic processor with a spike train sliding multicasting mechanism (STSM) and neuron merging mechanism (NMM) to support unconstrained fan-out and flexible fan-in of neurons. Unicorn supports 36K neurons and 45M synapses and thus supports a variety of neuromorphic applications. The peak performance and energy efficiency of Unicorn reach 36TSOPS and 424GSOPS/W respectively. Experimental results show that Unicorn can achieve 2×-5.5× energy reduction over the state-of-the-art neuromorphic processor when running an SNN with a relatively large fan-out and fan-in.
Lei Wang 0011, Yao Wang 0002, LingHui Peng, Xun Xiao, Weixia Xu 0001
DAC4
2022 An Event Based Gesture Recognition System Using a Liquid State Machine Accelerator
abstract
In this paper, we design a spiking neural network (SNN) accelerator based on the Liquid State Machine (LSM) which is more lightweight and bionic. In this accelerator, 512 leaky integrate-and-fire (LIF) neurons with configurable biological parameters are integrated. For the sparsity of computation and memory of the LSM, we use zero-skipping and weight compression to maximize the performance. The quantized 4-bit model deployed on the accelerator can achieve a classification accuracy of 97.42% on the DVS128 gesture dataset. We implement the accelerator on FPGA. Results indicate that its end-to-end average inference latency is 3.97 ms, which is 26 times better than the gesture recognition system based on TrueNorth.
Xun Xiao, Ziyang Kang, LingHui Peng
ACM Great Lakes Symposium on VLSI7
2021 A Novel Ring-based Small-World NoC for Neuromorphic Processor
abstract
Neuromorphic computing has shown promise in metrics such as power consumption and parallelism over existing computer systems, which essentially promote the development of neuromorphic processors in recent years. In order to properly place the increasing number of neuron cores and support the inter-core communication, Network-on-Chip (NoC) is widely used in the design of neuromorphic processors. Mesh has historically been used for multi-core NoCs, however, in neuromorphic chips, computation cores are relatively small, mesh-based SNN with high resource occupation limited the peak performance and energy efficiency. Moreover, one of the most significant findings in the neuroscience is that human brain exhibits small-world effect which is originated from the social network, inspired by that, we proposed a composite architecture for neuromorphic processor called the ring-based small-world NoC. Correspondingly, we proposed a routing algorithm to by generating a specific-application routing table in the pre-processing stage. We evaluated the performance such as delay, energy and resource utilization comprehensively based on three spike-based datasets (FSDD, NMNIST and N-TIDIGITS). The experimental results show that the average packet latency and the resource utilization of the proposed network is reduced by up to 18% and 35%, compared to a regular mesh network.
Yuchen Qiu, LingHui Peng, Ziyang Kang, Lei Wang 0011
ASAP3