VLDB 2026 Research / reviewers in the wild / expert
Ziyang Kang
dblp:249/7664
· DBLP profile ↗
24ranked-venue papers
8as first author
17since 2021 · last 2026
0000-0001-8798-365XORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 18 · 7 first-author · 12 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-author · 2 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | NPEva: An NoC-Based Neuromorphic Processors Performance Evaluation Framework for Benchmarking Spiking Neural NetworksabstractSpiking Neural Network (SNN) applications place diverse demands on neuromorphic processors’ computational and communication capabilities. Specifically, the sparse parallel computation of SNNs requires storage systems capable of extensive parallel data access and processing. Additionally, the time-dependent nature of SNN computations demands a communication framework to manage spike transmission and synchronization. Addressing these requires continuous early-stage evaluation of potential designs, which can be time-intensive. To this end, this paper introduces a performance evaluation model for NoC-based neuromorphic processors, named NPEva, which includes a Computation Model (CpMo) and a Communication Model (CoMo). This model facilitates rapid, high-dimensional exploration of design spaces across a wide range of microarchitectural parameters. CpMo quickly assesses computation-related latency, power consumption, and area by extracting relevant parameters from hierarchical storage organizations and neuron computation models. CoMo employs communication upper-bound latency analysis to evaluate packet latencies to reduce redundant synchronization time. Comparisons with TrueNorth’s publicly released data reveal that the model’s evaluation results are within an 8% error margin. The CoMo offers enhanced precision in estimating latency upper-bounds compared to other models like worst-contention delay (WCD) and worst-case traversal time (WCTT), especially when varying the number of VCs. With VC settings of 2, 4, 8, and 16, CoMo reduced latency by 0.55× to 2.53× compared to WCTT and by 0.83× to 25.57× compared to WCD. Additionally, using the NPEva model for exploring the design space of Liquid State Machines (LSM) networks has shown that optimized hardware can improve cost efficiency by 2.94× with only a 1% increase in latency. Ziyang Kang, Lei Wang 0011, De Ma, Gang Pan 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2025 | An NoC-Based Latency Upper-Bound Model for Reducing Timestep Length in SNN CommunicationabstractSpiking Neural Networks (SNNs) with real-time demands are deployed on Network-on-Chip (NoC)-based neuromorphic processors for specific tasks. While meeting real-time constraints for spiking data streams is crucial, overly long timesteps (e.g., 1ms) result in idle neuron cores and routers, reducing efficiency. This paper proposes a communication performance model using a recursive calculation method to assess the worst-case upper-bound latency. Integrated with the NoC router microarchitecture, the model analyzes the behavior of spiking streams during communication. It effectively balances timestep length to ensure real-time communication while minimizing idle time. Empirical results show that our model outperforms worst-contention delay (WCD) and worst-case traversal time (WCTT) models, achieving latency reductions of 0.55× to 2.53× compared to WCTT and 0.83× to 25.57× compared to WCD across various spiking datasets and virtual channels. Encouragingly, with only a 1% accuracy reduction, latency was reduced by 18%. Ziyang Kang, Lei Wang 0011, De Ma, Gang Pan 0001 |
ISCAS | 1 |
| 2025 | FCG: High-Throughput JPEG Heterogeneous Inference with Hybrid Parallel Pipeline on Mobile DevicesabstractWith the increasing popularity of image and video analysis on mobile devices, high-throughput image inference has become essential. However, current mobile deep learning frameworks face key bottlenecks: high computational load in JPEG image recognition and low processor efficiency, which limit overall image processing throughput. To address these issues, this paper proposes the FCG framework (Frequency Domain model for CPU and GPU), a mobile JPEG inference framework based on frequency domain data and a hybrid parallel architecture that enables high-throughput inference for JPEG-encoded images on mobile devices. FCG decouples JPEG decoding from model inference by discarding the traditional RGB decoding process and retaining only the Huffman decoding. This decoding step is further accelerated through multi-core processing, significantly reducing the computational burden and latency during preprocessing. In light of the characteristics of frequency domain data and the heterogeneous CPU/GPU processors on mobile devices, FCG reconstructs the deep learning model to ensure recognition accuracy while optimizing resource utilization. By effectively allocating tasks and combining parallel and sequential execution, FCG optimizes processor resource utilization to achieve high throughput and low latency. FCG outperforms the state-of-the-art NN-Stretch by reducing latency by 36%. It also achieves significant throughput improvements—3.6x, 3.3x, and 2.8x—on CPU, GPU, and CPU+GPU configurations, respectively, compared to sequential inference systems. Additionally, FCG reduces power consumption by 56%, 35%, and 43% in these configurations. Youbo Mao, Ziyang Kang, Jiyao Chen, Zenglin Yang, Zhijun Li 0002 |
ACM Multimedia | 2 |
| 2025 | HetSub: A Heterogeneous Multi-NoC With Reconfigurable Long-Range Links for Neuromorphic Systems
Youneng Hu, Xiaofei Jin, Ziyang Kang, De Ma, Gang Pan 0001 |
IEEE Trans. Very Large Scale Integr. Syst. | 4 |
| 2024 | LSM-Based Hotspot Prediction and Hotspot-Aware Routing in NoC-Based Neuromorphic ProcessorabstractThe traffic patterns of spiking neural networks (SNNs) exhibit high variability and stochastic, leading to the emergence of elevated traffic hotspots on the network-on-chip (NoC)-based neuromorphic processors. Predicting the occurrence of hotspots remains one of the most challenging issues in NoC design. This article presents the first attempt toward traffic hotspot prediction by utilizing liquid state machine (HP-LSM). The predictor extracts essential information reflecting the current state of the NoC to predict potential routing hotspots in the subsequent time step. Furthermore, we designed the hardware architecture for HP-LSM, which incorporates leaky-integrate-and-fire (LIF) neurons with configurable biological parameters. Meanwhile, we introduce a novel hotspot-aware path-based multicast (HaPM) routing algorithm that utilizes advanced knowledge acquired from HP-LSM to guide packet routing throughout the network, aiming to improve the performance of NoC. Results indicate that the HP-LSM can forecast hotspot formation with an accuracy up to 89.36% and 90.19% for two spiking-based datasets, respectively. The hardware experiment results demonstrate a 92.03% reduction in the average execution time of zero skipping compared with nonzero skipping. Moreover, the HP-LSM exhibits a reduction of up to 79.30% in the number of neurons compared with other related SNN predictor models. The experiments reveal a reduction of 73.67% and 53.42% in the average length of the multicast path when compared with dual-path (DP) or multipath (MP) multicast routing. The HaPM demonstrates improved performance in terms of average latency and throughput compared with DP, MP, and path-based multicast (PbM) multicast routing. Ziyang Kang, Xun Xiao, Lei Wang 0011, De Ma, Gang Pan 0001 |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2023 | Work-in-Progress: CLERR: A High-performance Cross-layer Method for Eliminating Rendering Redundancy in AndroidabstractRendering redundancy consumes a lot of computing resources in mobile devices. Eliminating redundancy can effectively improve system energy efficiency. However, as the premise of redundancy elimination, the existing redundancy identification methods affect the final performance because of the excessive cost. In this article, we propose the CLERR: a high-performance cross-layer method for eliminating rendering redundancy. CLERR decomposes the redundancy detection process into two collaborative steps, thereby reducing detection overhead. The proposed method is compatible with the mainstream Android 12 system. Experimental results indicate that the method can reduce frame drop rates by 14.5% and save SOC energy by 5.1%. Shixiong Huang, Nanxuan Ye, Xing Gao 0004, Ziyang Kang |
EMSOFT | 4 |
| 2023 | Path-Based Multicast Routing for Network-on-Chip of the Neuromorphic Processor
Ziyang Kang, Shi-Ying Wang, Lianhua Qu |
J. Comput. Sci. Technol. | 1 |
| 2022 | An Event Based Gesture Recognition System Using a Liquid State Machine AcceleratorabstractIn this paper, we design a spiking neural network (SNN) accelerator based on the Liquid State Machine (LSM) which is more lightweight and bionic. In this accelerator, 512 leaky integrate-and-fire (LIF) neurons with configurable biological parameters are integrated. For the sparsity of computation and memory of the LSM, we use zero-skipping and weight compression to maximize the performance. The quantized 4-bit model deployed on the accelerator can achieve a classification accuracy of 97.42% on the DVS128 gesture dataset. We implement the accelerator on FPGA. Results indicate that its end-to-end average inference latency is 3.97 ms, which is 26 times better than the gesture recognition system based on TrueNorth. Xun Xiao, Ziyang Kang, LingHui Peng |
ACM Great Lakes Symposium on VLSI | 5 |
| 2022 | Dynamic Vision Sensor Based Gesture Recognition Using Liquid State Machine
Xun Xiao, Lei Wang 0011, Lianhua Qu, Shasha Guo 0001, Yao Wang 0002, Ziyang Kang |
ICANN (3) | 7 |
| 2022 | A Spatio-Temporal Event Data Augmentation Method for Dynamic Vision Sensor
Xun Xiao, Ziyang Kang, Shasha Guo 0001, Lei Wang 0011 |
ICONIP (6) | 3 |
| 2022 | Hotspot Prediction of Network-on-Chip for Neuromorphic Processor with Liquid State MachineabstractThe traffic patterns of Spiking Neural Networks (SNNs) are highly varying and unpredictable, which will cause elevated traffic hotspots on the Network-on-Chip (NoC). How to predict the occurrence of hotspots remains one of the most challenging issues for NoC design. This work presents the first attempt toward utilizing a Liquid State Machine (LSM) in the prediction of traffic hotspot on routers. We also adopt the heuristic algorithm to search the hyperparameter of the LSM to improve the performance. Results indicate that the predictor can forecast hotspot formation with an accuracy up to 89.36% and 90.19% for two spiking-based datasets, respectively. Ziyang Kang |
ISCAS | 1 |
| 2022 | Hardware-aware liquid state machine generation for 2D/3D Network-on-Chip platforms
Ziyang Kang, Lei Wang 0011, Lianhua Qu, Weixia Xu 0001 |
J. Syst. Archit. | 1 |
| 2022 | LSMCore: A 69k-Synapse/mm2 Single-Core Digital Neuromorphic Processor for Liquid State MachineabstractNeuromorphic processors have gained momentum recently due to their high energy efficiency in artificial intelligence applications compared to DNN accelerators. Most neuromorphic processors are executing SNNs (Spiking Neural Networks). Liquid State Machine (LSM), as the spiking version of reservoir computing, shows advantages and great potential in image classification, speech recognition, language translation, etc.. Comparing with other SNN models, LSM has the characteristics of easy to train and low resource utilization, which is suitable for low-power and resource-constrained edge computing scenarios. In this paper, we propose a novel design of a neuromorphic processor, LSMCore, aiming at LSM acceleration. LSMCore supports both training and inference of LSM. It consists of 256 input neurons, 1024 liquid neurons, and 1.31M synapses. Besides, multiple optimization techniques, including weight quantization for reducing storage, zero-skipping for decreasing dynamic sparsity, and mini-batch training are adopted in this processor. The experimental results show that the frequency of LSMCore achieves 400 MHz, the power is 4.9W and the area is 18.49 mm2with a 40nm library. Comparing with the baseline, LSMCore achieves up to$80.7\times $($49.6\times $),$91.3\times $($56.3\times $), and$83.1\times $($56.8\times $) speedup on MNIST, N-MNIST, and Free Spoken Digital Dataset (FSDD) respectively for training (inference), while the accuracy of LSMCore on these three datasets are 96.8%, 97.6%, and 90% respectively. Lei Wang 0011, Shasha Guo 0001, Lianhua Qu, Ziyang Kang, Weixia Xu 0001 |
IEEE Trans. Circuits Syst. I Regul. Pap. | 6 |
| 2021 | A Novel Ring-based Small-World NoC for Neuromorphic ProcessorabstractNeuromorphic computing has shown promise in metrics such as power consumption and parallelism over existing computer systems, which essentially promote the development of neuromorphic processors in recent years. In order to properly place the increasing number of neuron cores and support the inter-core communication, Network-on-Chip (NoC) is widely used in the design of neuromorphic processors. Mesh has historically been used for multi-core NoCs, however, in neuromorphic chips, computation cores are relatively small, mesh-based SNN with high resource occupation limited the peak performance and energy efficiency. Moreover, one of the most significant findings in the neuroscience is that human brain exhibits small-world effect which is originated from the social network, inspired by that, we proposed a composite architecture for neuromorphic processor called the ring-based small-world NoC. Correspondingly, we proposed a routing algorithm to by generating a specific-application routing table in the pre-processing stage. We evaluated the performance such as delay, energy and resource utilization comprehensively based on three spike-based datasets (FSDD, NMNIST and N-TIDIGITS). The experimental results show that the average packet latency and the resource utilization of the proposed network is reduced by up to 18% and 35%, compared to a regular mesh network. Yuchen Qiu, LingHui Peng, Ziyang Kang, Lei Wang 0011 |
ASAP | 5 |
| 2021 | A Hardware Aware Liquid State Machine Generation FrameworkabstractThe liquid state machine (LSM) is a kind of spiking neural network (SNN) that usually is mapped to an NoC-based neuromorphic processor to perform tasks such as classification. The creation of these LSM models does not consider the structure of Network on Chip (NoC) which resulting in heavy communication pressure on the NoC. In this paper, we propose a hardware aware LSM network generation framework. By keeping the communication between neurons within cores as much as possible, this framework could reduce the communication overheads between cores effectively. The experimental results show that the LSM model produced by our framework could achieve state-of-art accuracy and is hardware-friendly. Compared with the mapping method, the synapses in our LSM is reduced by 94.14%, the total packets in NoC is reduced by 78.3%, the maximum transmission latency is reduced by 97.5%, the average transmission latency is reduced by 54%, the throughput increased 2.8x. Ziyang Kang, Lei Wang 0011, Lianhua Qu |
ISCAS | 2 |
| 2021 | HashHeat: A hashing-based spatiotemporal filter for dynamic vision sensor
Shasha Guo 0001, Ziyang Kang, Lei Wang 0011, Limeng Zhang, Weixia Xu 0001 |
Integr. | 2 |
| 2021 | A multi-objective LSM/NoC architecture co-design framework
Shuo Tian, Ziyang Kang, Lianhua Qu, Lei Wang 0011, Weixia Xu 0001 |
J. Syst. Archit. | 3 |
| 2020 | HashHeat: An O(C) Complexity Hashing-based Filter for Dynamic Vision SensorabstractNeuromorphic event-based dynamic vision sensors (DVS) have much faster sampling rates and a higher dynamic range than frame-based imagers. However, they are sensitive to background activity (BA) events which are unwanted. We propose HashHeat, a hashing-based BA filter with O(C) complexity. It is the first spatiotemporal filter that doesn't scale with the DVS output size N and doesn't store the 32-bits timestamps. HashHeat consumes 100x less memory and increases the signal to noise ratio by 15x compared to previous designs. Shasha Guo 0001, Ziyang Kang, Lei Wang 0011, Weixia Xu 0001 |
ASP-DAC | 2 |
| 2020 | Application-specific network-on-chip design space exploration framework for neuromorphic processorabstractNeuromorphic processors can support the design of various Spiking Neural Networks (SNN) to deal with different tasks, such as recognition and tracking. Neuromorphic processors use Network-on-Chip (NoC) to support communication between neurons in SNN. The different SNN has different communication traffic patterns. It will pose the different challenges of the NoC designing. A reasonable NoC architecture can improve the overall performance such as lower latency of the processor. Hence, it is critical to implement the exploration of NoC architecture design for neuromorphic processors. Ziyang Kang, Lei Wang 0011, Lianhua Qu, Weixia Xu 0001 |
CF | 1 |
| 2020 | SNEAP: A Fast and Efficient Toolchain for Mapping Large-Scale Spiking Neural Network onto NoC-based Neuromorphic PlatformabstractSpiking neural network (SNN), as the third generation of artificial neural networks, has been widely adopted in vision and audio tasks. Nowadays, many neuromorphic platforms support SNN simulation and adopt Network-on-Chips (NoC) architecture for multi-cores interconnection. However, a large volume and run-time communication on the interconnection has a significant effect on performance of the platform. In this paper, we propose a toolchain called SNEAP (Spiking NEural network mAPping toolchain) for mapping SNNs to neuromorphic platforms with multi-cores, which aims to reduce the energy and latency brought by spike communication on the interconnection. Shasha Guo 0001, Limeng Zhang, Ziyang Kang, Lei Wang 0011, Weixia Xu 0001 |
ACM Great Lakes Symposium on VLSI | 4 |
| 2020 | CompressedCache: Enabling Storage Compression on Neuromorphic Processor for Liquid State Machine
Lianhua Qu, Ziyang Kang, Lei Wang 0011, Weixia Xu 0001 |
NPC | 4 |
| 2020 | SIES: A Novel Implementation of Spiking Convolutional Neural Network Inference Engine on Field-Programmable Gate Array
Shuquan Wang, Lei Wang 0011, Yu Deng 0001, Shasha Guo 0001, Ziyang Kang, Yu-Feng Guo, Weixia Xu 0001 |
J. Comput. Sci. Technol. | 6 |
| 2020 | ASIE: An Asynchronous SNN Inference Engine for AER Events ProcessingabstractNeuromorphic computing based on spiking neural network (SNN) shows good energy-efficiency. However, it is inefficient for SNN to perform the convolution based on frame. It may contain a lot of redundant information in the frame. The output of Dynamic Vision Sensors (DVS) is a stream event based on Address Event Representation (AER). The asynchronous nature of AER events makes the event-based convolution reflect the characteristics of SNN low energy consumption. This article presents an SNN hardware inference engine based on an asynchronous Processing Element (PE) array with AER events as input. The engine uses a convolution algorithm based on AER events. This design also uses distributed storage in the PE array to store the state of neurons to reduce the cost of memory access. The experimental results show that the design can achieve a recognition accuracy of 98.0% for the MNIST AER dataset. The design can perform the reference process more efficiently in the case where the accuracy of the loss is negligible. During the filling and draining processes of the systolic array, the number of active PE units in our PE array is reduced and, thus, the average power consumption per PE unit is drastically decreased. Ziyang Kang, Lei Wang 0011, Shasha Guo 0001, Yu Deng 0001, Weixia Xu 0001 |
ACM J. Emerg. Technol. Comput. Syst. | 1 |
| 2019 | PRTSM: Hardware Data Arrangement Mechanisms for Convolutional Layer Computation on the Systolic Array
Shuquan Wang, Lei Wang 0011, Shuo Tian, Shasha Guo 0001, Ziyang Kang, Shuzheng Zhang, Weixia Xu 0001 |
NPC | 6 |