Chenglong Zou

dblp:68/7660 · DBLP profile ↗
← Back
8ranked-venue papers
1as first author
4since 2021 · last 2026
0000-0002-2571-3213ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 6 · 1 first-author · 3 since 2021Theory of computation · 1Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 PAICar: a prototype of an embodied neuromorphic intelligent robot platform
Mingkai Liu, Jingyi Zhong, Yi Zhong 0002, Zilin Wang 0001, Chenglong Zou, Xiaoxin Cui, Jian Cao 0002, Yuan Wang 0001
Sci. China Inf. Sci.7
2022 An Event-driven Spiking Neural Network Accelerator with On-chip Sparse Weight
abstract
Spiking neural networks (SNNs) have widely drew attention of recent research. With brain-spired dynamics and spike-based communication, SNN is supposed to be a more energy-efficient neural network than existing artificial neural network (ANN). To make better use of the temporal sparsity of spikes and spatial sparsity of weights in SNN, this paper presents a sparse SNN accelerator. It adopts a novel self-adaptive spike compressing and decompressing (SASCD) mechanism for different input spike sparsity, as well as on-chip compressed weight storage and processing. We implement the octa-core design on field programmable gate array (FPGA). The results demonstrate a peak performance of 35.84 GSOPs/s, which is equivalent to 358.4 GSOPs/s in dense SNN accelerators for 90% weight sparsity. For the single-layer perceptron model in rate coding implemented on the hardware, SASCD reduces the time step intervals from 2.15 $\mu$ s to 0.55 $\mu$ s.
Yisong Kuang, Xiaoxin Cui, Chenglong Zou, Yi Zhong 0002, Zhenhui Dai, Zilin Wang 0001, Kefei Liu 0002, Dunshan Yu, Yuan Wang 0001
ISCAS3
2022 ESSA: Design of a Programmable Efficient Sparse Spiking Neural Network Accelerator
abstract
Spiking neural networks (SNNs) have been witnessing the developing trends to reduce the model size and improve the hardware efficiency for area- and energy-based applications, which are processed by model pruning and data compressions. However, it is challenging to exploit the unstructured sparsity of SNNs for the dense neuromorphic processors. In this article, we present an efficient sparse SNN accelerator (ESSA), which leverages both the temporal sparsity of spike events and the spatial sparsity of weights in SNN inference. It provides both the compressed weights for sparse SNNs and the uncompressed weights for compact SNNs. The self-adaptive spike compression is proposed for sparse spike scenarios, leading to the improvement of throughput by$3.2\times $. ESSA executes a flexible fan-in–fan-out tradeoff by using combinable dendrites, which overcomes the fan-in limitation in neuromorphic systems. Furthermore, a low-latency intrachip spike multicast method is adopted to reduce the resource overhead. Implemented on the Xilinx Kintex Ultrascale field-programmable gate array (FPGA), ESSA achieves an equivalent performance of 253.1 GSOP/s and an energy efficiency of 32.1 GSOP/W for 75% weight sparsity at 140 MHz. The implementation of a four-layer fully connected SNN is expected to perform$2.6~\mu \text{s}$per time step and the energy consumption is$14.6~\mu \text{J}$. Our results demonstrate that ESSA outperforms several state-of-the-art application-specific integrated circuit (ASIC) or FPGA neuromorphic processors.
Yisong Kuang, Xiaoxin Cui, Zilin Wang 0001, Chenglong Zou, Yi Zhong 0002, Kefei Liu 0002, Zhenhui Dai, Dunshan Yu, Yuan Wang 0001, Ru Huang 0001
IEEE Trans. Very Large Scale Integr. Syst.4
2021 A 28-nm 0.34-pJ/SOP Spike-Based Neuromorphic Processor for Efficient Artificial Neural Network Implementations
abstract
Neuromorphic hardware platforms inspired by human brain have emerged as novel non von Neumann computing architectures. They were proved excellent platforms for spiking neural network (SNN) implementations. However, implementing artificial neural networks (ANNs) on existing neuromorphic hardware platforms is still a daunting task because of critical limitations on coding scheme, maximum of fan-in, and highest weight precision in them. In this paper, we introduce a neuromorphic processor developed for various neural networks implementations including ANNs and SNNs. We employ spatio-temporal coding scheme based on spike events. By combining low-precision dendrites, the chip can implement weight precision between 1 bit and 8 bits and scalable fan-in. The 3.66-mm2chip fabricated in 28-nm CMOS with a maximum fan-in of 72 K per neuron demonstrates unprecedented compatibility with ANN applications compared to previously-proposed neuromorphic chips.
Yisong Kuang, Xiaoxin Cui, Yi Zhong 0002, Kefei Liu 0002, Chenglong Zou, Zhenhui Dai, Dunshan Yu, Yuan Wang 0001, Ru Huang 0001
ISCAS5
2020 A Novel Conversion Method for Spiking Neural Network using Median Quantization
abstract
Artificial Neural Networks (ANNs) have achieved great success in the field of computer vision and language understanding. However, it is difficult to deploy these deep learning models on mobile devices because of its massive energy consumption and memory occupation. For another way, highly inspired from biological brain, spiking neural networks (SNNs), are often referred to as the 3-th generation of neural network for its potential superiority in cognitive learning and energy efficiency. Nevertheless, training a deep SNN remains a big challenge. In this paper, we propose a quantized training algorithm for ANNs to minimize spike approximation error, and provide two (temporally or spatially) rate-based conversion methods for SNNs, both of which can be easily mapped to specific neuromorphic platforms. Besides, this novel method can be generalized to various network architectures and adapted to dynamic quantization demand. Experimental results on MNIST and CIFAR-10 dataset demonstrate that the proposed deep spiking neural networks yield the state-of-the-art classification accuracy and need much less operations compared with their ANN counterparts. Our source code will be available upon request for the academic purpose.
Chenglong Zou, Xiaoxin Cui, Jiexian Ge, Hanghang Ma, Xin'an Wang
ISCAS1
2018 BenchBox: A User-Driven Benchmarking Framework for Fat-Client Storage Systems
abstract
In many online storage services, end-users mainly interact with the system via “fat” storage clients that integrate complex functionality. This means that to obtain a complete performance evaluation of one of such systems we may need to generate workloads on the client side that reproduce the behavior of real users. Unfortunately, this remains as an open research challenge today. We present BenchBox: A distributed performance evaluation framework for fat-client storage systems. On the one hand, BenchBox can generate workloads directly in storage clients that mimic users exhibiting a certain behavior, namely, user stereotypes. To this end, the framework enables to plug-in workload models and feed them with compact recipes that capture the behavior of user stereotypes (e.g., storage activity, type of file contents, data sharing links). On the other hand, BenchBox provides researchers with management and monitoringfacilities to deploy experiments and analyze the performance of groups of storage clients. To demonstrate our framework, we equipped BenchBox with a 2-layer workload modelthat reproduces both the activity-e.g., types of operations, frequency- and data-e.g., file sizes, data types-of users in a Personal Cloud. We used this model to generate workloads based on user stereotypes that we identified in real traces (UbuntuOne). Our experiments with public providers show how distinct types of users impact on the performance and efficiency of Personal Clouds, which may guide their optimization.
Raúl Gracia Tinedo, Chenglong Zou, Marc Sánchez Artigas, Pedro García López
IEEE Trans. Parallel Distributed Syst.2
2017 COSY: An Energy-Efficient Hardware Architecture for Deep Convolutional Neural Networks Based on Systolic Array
abstract
Deep convolutional neural networks (CNNs) show extraordinary abilities in artificial intelligence applications, but their large scale of computation usually limits their uses on resource-constrained devices. For CNN's acceleration, exploiting the data reuse of CNNs is an effective way to reduce bandwidth and energy consumption. Row-stationary (RS) dataflow of Eyeriss is one of the most energy-efficient state-of-the-art hardware architectures, but has redundant storage usage and data access, so the data reuse has not been fully exploited. It also requires complex control and is intrinsically unable to skip over zero-valued inputs in timing. In this paper, we present COSY (CNN on Systolic Array), an energy-efficient hardware architecture based on the systolic array for CNNs. COSY adopts the method of systolic array to achieve the storage sharing between processing elements (PEs) in RS dataflow at the RF level, which reduces low-level energy consumption and on-chip storage. Multiple COSY arrays sharing the same storage can execute multiple 2-D convolutions in parallel, further increasing the data reuse in the low-level storage and improving throughput. To compare the energy consumption of COSY and Eyeriss running actual CNN models, we build a process-based energy consumption evaluation system according to the hardware storage hierarchy. The result shows that COSY can achieve an over 15% reduction in energy consumption under the same constraints, improving the theoretical Energy-Delay Product (EDP) and Energy-Delay Squared Product (ED2P) by 1.33× on average. In addition, we prove that COSY has the intrinsic ability for zero-skipping, which can further increase the improvements to 2.25× and 3.83× respectively.
Miren Tian, Mohan Ji, Chenglong Zou, Xin'an Wang, Bo Wang 0016
ICPADS5
2014 Quotient Complexity of Closed Languages
Janusz A. Brzozowski, Galina Jirásková, Chenglong Zou
Theory Comput. Syst.3