VLDB 2026 Research / reviewers in the wild / expert
Shanlin Xiao
dblp:190/8675
· DBLP profile ↗
25ranked-venue papers
4as first author
20since 2021 · last 2026
0000-0002-1250-8704ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 15 · 3 first-author · 12 since 2021Artificial intelligence and machine learning · 6 · 1 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 4 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Spiking-NeRF: Neural Graphics Acceleration With Spiking Feature Encoding for Edge 3D RenderingabstractNeural Radiance Fields (NeRF) have demonstrated remarkable potential for high-fidelity 3D scene reconstruction and rendering. However, achieving real-time performance remains a major challenge due to two critical bottlenecks: the high memory demand of multi-resolution hash encoding and the considerable computational cost of floating-point interpolation. To address these limitations, we propose Spiking-NeRF, a braininspired algorithm-hardware co-design framework. On the algorithm side, we introduce a spiking feature encoding scheme based on Integrate-and-Fire (IF) neurons, which transforms continuous voxel features into sparse spikes, reducing hash storage overhead by $\mathbf{7 5 \%}$. We further propose a global importance-based pruning strategy that compresses hash tables by $\mathbf{7 1. 3 \%}$ by removing lowaccessed entries. To reduce interpolation complexity, we design a hard-threshold weight discretization method that eliminates floating-point operations in favor of bitwise logic. On the hardware side, we accelerate critical stages of the NeRF pipeline by integrating a spike-skipping mechanism that dynamically bypasses hash entries, reducing memory traffic by 32.46%. We also co-optimize on-chip storage by leveraging access locality patterns across different resolution levels of the hash structure. Experimental results demonstrate that Spiking-NeRF achieves real-time rendering performance while maintaining high visual fidelity. Compared to edge GPUs, our design improves throughput by $111.2 \times$ and reduces power consumption by $41.67 \times$. Against SOTA NeRF accelerators, Spiking-NeRF achieves up to $2.48 \times$ higher throughput and $5.45 \times$ lower energy usage, underscoring the potential of spike-based computing for next-generation lowpower neural graphics systems. Jianzhen Gao, Wei Liu 0118, Hengyi Zhou, Zhiyi Yu, Shanlin Xiao |
ASP-DAC | 6 |
| 2026 | An Algorithm-Hardware Co-Design for Efficient and Robust Spiking Neural Networks via SparsityabstractSNN deployment faces a dilemma: rate codes are power-hungry while temporal codes lack noise resilience. This paper proposes an SNN algorithm-hardware co-design, which uses sparse coding and a zero-skipping accelerator to alleviate the rate-temporal trade-off. The design reduces network spike count by 88% compared to rate coding while enhancing fault tolerance. Benchmarked against state-of-the-art rate-coding and temporal-coding accelerators, the prototype saves 88% and 89% energy, achieves $4.5 \times$ and $26.8 \times$ higher throughput, and uses 82% fewer LUTs, enabling efficient and robust edge inference. Wei Liu 0118, Yinsheng Chen, Jilong Luo, Yusa Wang, Zhiyi Yu, Shanlin Xiao |
ASP-DAC | 6 |
| 2026 | SpikeVPR: Energy-efficient visual place recognition via multi-scale spiking transformers
Hengyi Zhou, Yinsheng Chen, Jilong Luo, Jianzheng Gao, Wei Liu 0118, Shanlin Xiao |
Neurocomputing | 6 |
| 2026 | SynapseHD: A unified training framework for bridging spiking neural networks and hyperdimensional computing
Lingfeng Zhou, Huiyao Wang, Jinghai Wang, Zhiyi Yu, Shanlin Xiao |
Neurocomputing | 6 |
| 2025 | SCSC: Leveraging Sparsity and Fault-Tolerance for Energy-Efficient Spiking Neural NetworksabstractSpiking neural networks (SNNs) are more energy-efficient for processing sparse spike signals and demonstrate better fault tolerance compared to artificial neural networks (ANNs). In neuromorphic chips, synaptic weight access and neuron computation operations constitute 75%-95% of the chip's energy consumption. Therefore, our primary strategy to achieve highly energy-efficient SNNs is to enhance network sparsity while leveraging SNNs' high fault tolerance to reduce both weight access and neuron computation energy. The coding module is an essential component of SNNs, responsible for encoding non-spiking inputs into spike trains. However, previous coding schemes often exhibit poor sparsity or fault tolerance performance. Thus, we propose a novel coding scheme for SNNs: spiking convolutional sparse coding (SCSC). SCSC utilizes convolutional kernels as dictionaries and achieves sparsity through neural layers. Additionally, dynamic firing thresholds in neural layers balance sparsity with network performance and fault tolerance. The experimental results indicate that SCSC can increase network sparsity by 10%-20% and achieves higher accuracy than baseline networks when dealing with disturbances. Furthermore, we utilize approximate DRAM to store synaptic weights and selectively deactivate specific neuronal computing modules. With only a 1% decrease in accuracy, SCSC can reduce synaptic weight access energy by 29% and neuronal computing energy by 49%. Wei Liu 0118, Jinghai Wang, Zhiyi Yu, Shanlin Xiao |
ASP-DAC | 7 |
| 2025 | Towards In-Situ Neuromorphic Computing Architecture for Event Stream Super-ResolutionabstractEvent-based cameras, with their unique event stream representation, effectively mitigate motion blur in highspeed, high-exposure scenarios but suffer from low spatial resolution. To address this, we propose a super-resolution hardware accelerator for event streams based on Spiking Neural Networks (SNNs). In terms of network architecture, we incorporate hardware-friendly algorithmic designs by simplifying neuron models and optimizing convolution operations. On the hardware side, the design adopts a hierarchical structure featuring a highly parallel computational array. Additionally, by proposing a Kernel-Channel-Timestamp-Row (KCTR) dataflow and dual-pipeline structure, the design achieves in-situ computing, eliminating intermediate storage within layers and significantly reducing inter-layer spike storage. Evaluations on the N-MNIST and ASL-DVS datasets demonstrate root mean square errors (RMSE) of $\mathbf{1. 2 9 6}$ and $\mathbf{0. 1 2 1}$ for reconstructed super-resolution event streams. In downstream applications, the classification accuracies reach 98.84% and 99.73%, respectively. The proposed accelerator, designed using a 28 nm CMOS process, improves reconstruction speed by 95.6% compared to a GPU, operates at 500 MHz, and consumes only 0.546 pJ per synaptic operation. Yihe Yu, Wei Liu 0118, Jinghai Wang, Zhiyi Yu, Shanlin Xiao |
DAC | 8 |
| 2025 | Fine-Grained Recognition of Arteriovenous Fistula Stenosis Using Blood Flow Sounds: An Animal Model-Based Dataset and a Frequency-Aware Decoupling Network
Shanlin Xiao, Yangyi Zhou, Jincen Wang, Yuan Zong |
ICANN (4) | 1 |
| 2025 | Unsupervised Motion-Robust Self-Distillation Framework for Remote Physiological MeasurementabstractRemote photoplethysmography (rPPG) holds great potential in medical surveillance. However, head movements commonly encountered in real-world scenarios often degrade physiological estimation performance, particularly for unsupervised learning methods based on physiological frequency band priors, which usually struggle to detect dynamic signals occurring within this band, thereby limiting their performance ceiling. In this paper, we propose a novel strategy to endow unsupervised learning methods with motion robustness. Specifically, we introduce a simple motion simulation technique, Sliding Crop, to incorporate dynamic signals. Based on this, we develop an unsupervised motion-robust self-distillation framework (UMoRo) with the existing unsupervised learning method, where the model leverages its own high-quality mappings of simple samples as pseudo-labels to guide the learning process of suppressing simulated motion artifacts in challenging samples, thus enhancing motion robustness in real-world movements. Experimental results on three public datasets show that our method achieves superior or competitive performance compared to state-of-the-art supervised methods, demonstrating outstanding motion robustness. Anbang Liu, Shanlin Xiao, Wenming Zheng |
ICASSP | 2 |
| 2025 | Faster-SNN: Towards Faster and Better Spiking Neural Networks with Hybrid Neural CodingabstractInspired by the heterogeneity of the brain, hybrid neural coding in SNN models has garnered increasing attention from researchers. However, most prior research relies on the ANN2SNN conversion method, which results in large time steps and decreased energy efficiency. To overcome these limitations, we propose a faster SNN model (Faster-SNN) based on hybrid neural coding and direct training. Faster-SNN assigns different coding schemes to the input layer, hidden layer, and output layer to achieve hybrid coding. The input layer uses temporal & spatial attention coding (TSAC), which incorporates a spatio-temporal attention mechanism to enhance spatio-temporal information processing. In addition, the hidden layer employs an optimized Burst-LIF neuron to implement burst coding, effectively leveraging the residual information in the membrane potential to improve information transfer efficiency. Finally, the output layer uses TTFS coding to ensure accurate and rapid decision-making. Experimental results demonstrate that our model achieves high-accuracy inference with extremely low latency through the use of hybrid neural coding and direct training methods. Yinsheng Chen, Jilong Luo, Zhiyi Yu, Shanlin Xiao |
ICME | 4 |
| 2025 | Stair-LIF: Boosting the Representation of Spiking Neural Networks with Learnable Incremental Multi-Threshold NeuronsabstractSpiking neural networks (SNNs) have shown remarkable potential in processing spatio-temporal data by mimicking biological neuronal mechanisms and achieving low computational costs. However, previous SNNs often rely on neuron models with fixed and single threshold voltages and binary spikes across layers during training, which limits their capacity for accurate information representation and reduces their biological plausibility. Inspired by the diversity of neuronal behaviors in different brain regions, we propose a novel neuron called Stair-LIF, which introduces learnable incremental multi-threshold mechanisms to enhance neuronal representational capacity and utilizes multi-spike firing to improve the precision of information transmission. Furthermore, we propose a channel-wise parameterization method to expand representational capacity among Stair-LIF. Experimental results on static datasets (CIFAR-10, CIFAR-100) and neuromorphic dynamic datasets (CIFAR10-DVS and DVS128 Gesture) demonstrate that the Starir-LIF neuron achieves state-of-the-art performance. Jilong Luo, Yinsheng Chen, Jinghai Wang, Zhiyi Yu, Shanlin Xiao |
ICME | 6 |
| 2025 | A Hybrid Stochastic-Binary Computing Batch Normalization Engine for Low-Power On-Chip Learning Spiking Neural NetworksabstractBatch normalization (BN) has proven to be a critical component in speeding up the training of deep spiking neural networks in deep learning. However, conventional BN implementations face significant challenges in terms of excessive off-chip memory bandwidth requirements and complex circuit designs, hindering their applicability for on-chip training in spiking neural networks (SNNs). This article introduces a novel hybrid stochastic-binary computing BN engine (HBN) that strikes an optimal balance between computational efficiency and hardware resource utilization, enabling efficient on-chip learning for SNNs. While conventional binary-mode BN engines offer temporal efficiency, they demand substantial hardware resources. In contrast, stochastic computing (SC)-based BN approaches reduce hardware overhead but introduce latency penalties and necessitate additional random number generation (RNG) circuits. To overcome these limitations, we propose a hybrid architecture that seamlessly integrates binary and stochastic computing (SC) paradigms. Our co-designed methodology effectively balances computational latency and hardware footprint. This is achieved by a rounding-free SC multiplier unified with binary-circuit map ping, which eliminates latency and RNG overheads. Extensive validation across both static image datasets and neuromorphic datasets demonstrates that HBN maintains algorithmic fidelity while achieving unprecedented computational efficiency. Simulation results reveal 98.7% reduction in floating-point operations (FLOPs), 98.5% latency improvement, and 98.2% energy consumption reduction compared with conventional BN implementations. FPGA implementation on the ZCU102 platform demonstrates practical hardware advantages, including 74.9% reduction in look-up table (LUT) utilization, 83.6% decrease in flip-flop (FF) count, and 13.7% reduction in block RAM (BRAM) allocation. Notably, the design achieves 63.7% power reduction compared with state-of-the-art implementations while maintaining complete DSP-free operation. Wei Liu 0118, Zhiyi Yu, Shanlin Xiao |
IEEE Trans. Very Large Scale Integr. Syst. | 4 |
| 2025 | ASNA-Flow: An Efficient Asynchronous Neuromorphic Accelerator for Real-Time Event-Based Optical FlowabstractOptical flow estimation constitutes a fundamental computational challenge in computer vision, with critical applications object trajectory prediction, depth reconstruction, and autonomous navigation systems. The emergence of neuromorphic vision systems, integrating event-driven cameras with spiking neural networks (SNNs), has recently gained attention as a promising paradigm for edge deployment of optical flow estimation due to their advantages in ultralow power and resource efficiency. However, current neuromorphic computing platforms lack specialized architectures optimized for this problem domain. Existing implementations either prioritize configurable architectures at the expense of energy efficiency or employ intricate hardware control mechanisms to manage the asynchronous and sparse computing patterns inherent in SNNs. To address these limitations, we present ASNA-Flow, an event-driven asynchronous neuromorphic accelerator featuring a pioneering algorithm–hardware co-design framework specifically tailored for event-based optical flow estimation. Our methodology encompasses three key innovations: 1) a hardware-aware algorithm optimization that maintains computational fidelity while enhancing implementation efficiency; 2) systematic data pattern analysis to inform architectural decisions; and 3) novel exploitation of optical flow’s spatial locality characteristics to enable efficient sparse computing. Implemented in TSMC 28-nm CMOS technology, ASNA-Flow achieves real-time performance of 104 frames per second (FPS) with ultralow power consumption of 7.9 mW, demonstrating superior energy efficiency of 0.3 pJ per synaptic operation (SOP). This work establishes the first dedicated neuromorphic computing solution that simultaneously addresses the temporal sparsity, event-driven processing, and energy constraints inherent in optical flow estimation tasks. Jinghai Wang, Jilong Luo, Lingfeng Zhou, Zhiyi Yu, Shanlin Xiao |
IEEE Trans. Very Large Scale Integr. Syst. | 6 |
| 2025 | An End-to-End Bundled-Data Asynchronous Circuits Design Flow: From RTL to GDSabstractAsynchronous circuits with low power and robustness are revived in emerging applications such as the Internet of Things (IoT) and neuromorphic chips, thanks to clock-less and event-driven mechanisms. However, the lack of mature computer-aided design (CAD) tools for designing large-scale asynchronous circuits results in low design efficiency and high cost. This article proposes an end-to-end bundled-data (BD) asynchronous circuit design flow, which can facilitate building asynchronous circuits, even if the designer has little or no asynchronous circuit foundation. Three features that enable this are: 1) a lightweight circuit converter developed in Python can convert circuits from synchronous descriptions to corresponding asynchronous ones at register transfer level (RTL). Desynchronization flow helps designers maintain a “synchronization mentality” to construct asynchronous circuits; 2) a synchronization-like verification method is proposed for asynchronous circuits so that it can be functionally verified before synthesis. Avoids the risk of rework after logic defects are discovered during the synthesis and implementation, as asynchronous circuits often cannot be simulated until gate-level (GL) netlist generation; and 3) the whole implementation flow from RTL to graphic data system (GDS) is based on commercial electronic design automation (EDA) tools. Similar to the design flow of synchronous circuits, it helps designers implement asynchronous circuits with “synchronization habits.” Furthermore, to validate this methodology, two asynchronous processors were, respectively, implemented and evaluated in the TSMC 28-nm CMOS process. Compared to their synchronous counterparts, the general-purpose asynchronous RISC-V processor achieves 20.5% power savings. And the domain-specific asynchronous spiking neural network (SNN) accelerator achieves 58.46% power savings and$2.41\times $energy efficiency improvement at 70% input spike sparsity. Jinghai Wang, Shanlin Xiao, Jilong Luo, Lingfeng Zhou, Zhiyi Yu |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2024 | An Efficient Asynchronous Circuits Design Flow with Backward Delay Propagation ConstraintabstractAsynchronous circuits have recently become more popular in Internet of Things (IoT) and neural network chips because of their potential low power consumption. However, due to the lack of Electronic Design Automation (EDA) tools, the asynchronous circuits design efficiency remains low and faces challenges in large-scale applications. This paper proposes a new asynchronous circuits design flow using traditional EDA tools, and applies a new backward delay propagation constraint (BDPC) method. In this method, control paths and data paths are tightly coupled and analyzed together to improve the accuracy of static timing analysis. Compared to previous works, the proposed design flow and constraint method offer significant advantages in terms of accuracy and efficiency. To verify this flow, an asynchronous RISC-V processor was implemented on TSMC 65nm process. Compared to synchronous version, asynchronous processor achieves a power optimization of 17.4 % while main-taining the same speed and area. Lingfeng Zhou, Shanlin Xiao, Huiyao Wang, Jinghai Wang, Zeyang Xu, Zhiyi Yu |
DATE | 2 |
| 2024 | Sparsespikformer: A Co-Design Framework for Token and Weight Pruning in Spiking TransformerabstractAs the third-generation neural network, the Spiking Neural Network (SNN) has the advantages of low power consumption and high energy efficiency, making it suitable for implementation on edge devices. However, despite these advantages, SNN still faces accuracy limitations when compared to Artificial Neural Network (ANN). More recently, the most advanced SNN, Spikformer, combines the self-attention module from Transformer with SNN to achieve accuracy comparable to that of ANN. Additionally, to improve the final accuracy, the researchers adopt larger channel dimensions in MLP layers, leading to an increased number of redundant model parameters. To effectively decrease the computational complexity and weight parameters of the model, we explore the Lottery Ticket Hypothesis (LTH) and discover a very sparse (>90%) subnetwork that achieves comparable performance to the original network. Furthermore, we also design a lightweight token selector module, which can remove unimportant background information from images based on the average spike firing rate of neurons, selecting only essential foreground image tokens to participate in attention calculation. Experimental results demonstrate that our co-design framework can significantly reduce 90% model parameters and cut down Giga Floating-Point Operations (GFLOPs) by 20% while maintaining the accuracy of the original model. Shanlin Xiao, Zhiyi Yu |
ICASSP | 2 |
| 2024 | Better-Than-Worst-Case: A Frequency Adaptation Asynchronous RISC-V Core With Vector ExtensionabstractIn recent years, asynchronous circuits have become more popular in neural network chips and the Internet of Things (IoT) due to their potential advantages of low-power consumption and high performance. However, the existing design methods for asynchronous circuits are still constrained by critical paths, increasing power consumption and hindering the further improvement of performance. In this article, a fully digital design method for frequency adaptation asynchronous bundled-data (BD) circuits is proposed. The proposed method is straightforward, effective, widely applicable, and independent of asynchronous controllers. It allows to automatically work on different frequencies as required, which can improve performance and reduce power consumption, achieving better-than-worst-case. To verify the proposed method, an asynchronous RISC-V processor with vector acceleration extension is designed on both TSMC 65-nm process and field-programmable gate array (FPGA) platform. According to the postlayout simulation results, compared with its synchronous version, the asynchronous processor achieves a 10% speed improvement (from 227.3 to 250 MHz) with a 37% power reduction (from 135 to 85$\mu$W/MHz) under ideal conditions. Even under the worst conditions, the asynchronous processor achieves equivalent performance to the synchronous processor, while still reducing power consumption by 29% (from 133 to 95$\mu$W/MHz). On the FPGA platform, asynchronous processor also achieves higher speed while lower power consumption. Lingfeng Zhou, Shanlin Xiao, Huiyao Wang, Jinghai Wang, Zeyang Xu, Zhiyi Yu |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2024 | Toward Efficient Asynchronous Circuits Design Flow Using Backward Delay Propagation ConstraintabstractIn recent years, asynchronous circuits have gained attention in neural network chips and Internet of Things (IoT) due to their potential advantages of low power and high performance. However, design efficiency of asynchronous circuits remains low and faces challenges in large-scale applications because of the lack of electronic design automation (EDA) support. This article presents a new bundled-data (BD) asynchronous circuits’ design flow using traditional EDA tools, including a new backward delay propagation constraint (BDPC) method. In this method, control paths and data paths are analyzed together in a tightly coupled approach to improve the accuracy of static timing analysis (STA). Compared with other design flows, the proposed design flow and constraint method show significant advantages in aspects of STA accuracy, design efficiency, and design applicability, and solving the congestion issues of field-programmable gate array (FPGA) in a previous work. An asynchronous RISC-V processor was implemented to verify the method, with selective handshake technology to further reduce power. Compared with the synchronous processor, the asynchronous processor achieves a 17.4% power optimization on the TSMC 65-nm process and a 48.3% dynamic power savings on the FPGA while maintaining the same frequency and resource utilization. Lingfeng Zhou, Shanlin Xiao, Huiyao Wang, Jinghai Wang, Zeyang Xu, Zhiyi Yu |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2021 | 3D-VNPU: A Flexible Accelerator for 2D/3D CNNs on FPGAabstractThree-dimensional convolutional neural networks (3D CNNs) have proven to be outstanding in applications such as video analysis, 3-dimension geometric data, and 3-dimension medical image diagnosis. Compared to 2D CNNs, 3D CNNs require high computational complexity to get spatio-temporal features while Winograd algorithm can significantly reduce the amount of computation. Prior works based on 3D Winograd accelerators are only applied to stride-1 convolution, however, most of the popular 3D CNNs contain stride-2 convolution layers. In this paper, we propose a novel flexible Winograd-based decomposition method (FWDM) to apply the 3D Winograd to different strides convolution. Evaluation results show that FWDM reduces computational complexity by a factor of 3.2 for C3D, 2.9 for 3D ConvNet, and 2.6 for 3D ResNet-18. Furthermore, we design a flexible computing engine to stretch the use range of the decomposition method. Coupling FWDM and computing engine, a Winograd-based, 2D/3D CNNs compatible, highly efficient, and flexible accelerator (3D-VNPU) is proposed. Finally, we demonstrate the effectiveness of 3D-VNPU on FPGA platform (Xilinx ZCU102) and achieve 1.35TOPS for C3D, 1.2TOPS for 3D ResNet-18, and 1.1TOPS for VGG-16. DSP efficiency outperforms other CNNs accelerators 2.57~15.3x compared with prior works in FPGA of C3D. Compared to GPU and CPU, our accelerator achieves improvement up to 37.9x in performance relative to CPU and 11.8x in energy efficiency relative to GPU. Huipeng Deng, Jian Wang 0080, Huafeng Ye, Shanlin Xiao, Xiangyu Meng 0003, Zhiyi Yu |
FCCM | 4 |
| 2021 | High-parallelism Inception-like Spiking Neural Networks for Unsupervised Feature Learning
Mingyuan Meng, Lei Bi 0001, Jinman Kim, Shanlin Xiao, Zhiyi Yu |
Neurocomputing | 5 |
| 2021 | A Data-Driven Asynchronous Neural Network AcceleratorabstractDeep neural networks (DNNs) are revolutionizing machine learning, with unprecedented accuracy on many AI tasks. Energy-efficient neural acceleration is crucial in broadening DNN applications in cloud and mobile end devices. However, power-hungry clock networks limit the energy-efficiency of DNN accelerators. In this work, we propose a novel DNN hardware accelerator, called the asynchronous neural network processor (AsNNP). At the heart of AsNNP is a scalable hierarchy matrix multiply unit, with bit-serial processing elements working in parallel. It replaces the global clock networks with asynchronous handshake protocols to realize the synchronization and communication between each part, minimizing the dynamic power. Meanwhile, a fine-grain asynchronous pipeline based on weak-conditioned half-buffer (WCHB) is introduced to pipe successive computations in a data-driven manner, i.e., once data arrives computation begins, maximizing the throughput. These techniques enable AsNNP to work in a fully data-driven asynchronous communication fashion with optimized energy-efficiency. The proposed accelerator is implemented with quasi-delay-insensitive (QDI) clockless logic family and evaluated in a 65 nm process. Compared with the synchronous baseline, simulation results show that AsNNP offers 2.2× higher equivalent frequency and 1.59× lower power. Compared with state-of-the-art DNN accelerators, AsNNP shows 1.17×-4.97× energy-efficiency improvement. Shanlin Xiao, Weikun Liu, Junshu Lin, Zhiyi Yu |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2020 | SPA: Stochastic Probability Adjustment for System Balance of Unsupervised SNNsabstractSpiking neural networks (SNNs) receive widespread attention because of their low-power hardware characteristic and brain-like signal response mechanism, but currently, the performance of SNNs is still behind Artificial Neural Networks (ANNs). We build an information theory-inspired system called Stochastic Probability Adjustment (SPA) system to reduce this gap. The SPA maps the synapses and neurons of SNNs into a probability space where a neuron and all connected pre-synapses are represented by a cluster. The movement of synaptic transmitter between different clusters is modeled as a Brownian-like stochastic process in which the transmitter distribution is adaptive at different firing phases. We experimented with a wide range of existing unsupervised SNN architectures and achieved consistent performance improvements. The improvements in classification accuracy have reached 1.99% and 6.29% on the MNIST and EMNIST datasets respectively. Mingyuan Meng, Shanlin Xiao, Zhiyi Yu |
ICPR | 3 |
| 2020 | Spiking Inception Module for Multi-layer Unsupervised Spiking Neural NetworksabstractSpiking Neural Network (SNN), as a brain-inspired approach, is attracting attention due to its potential to produce ultra-high-energy-efficient hardware. Competitive learning based on Spike-Timing-Dependent Plasticity (STDP) is a popular method to train an unsupervised SNN. However, previous unsupervised SNNs trained through this method are limited to a shallow network with only one learnable layer and cannot achieve satisfactory results when compared with multi-layer SNNs. In this paper, we eased this limitation by: 1) We proposed a Spiking Inception (Sp-Inception) module, inspired by the Inception module in the Artificial Neural Network (ANN) literature. This module is trained through STDP-based competitive learning and outperforms the baseline modules on learning capability, learning efficiency, and robustness. 2)We proposed a Pooling-Reshape-Activate (PRA) layer to make the Sp-Inception module stackable. 3)We stacked multiple Sp-Inception modules to construct multilayer SNNs. Our algorithm outperforms the baseline algorithms on the hand-written digit classification task, and reaches state-of-the-art results on the MNIST dataset among the existing unsupervised SNNs. Mingyuan Meng, Shanlin Xiao, Zhiyi Yu |
IJCNN | 3 |
| 2020 | A Low-Cost and High-Throughput NoC-Aware Chip-to-Chip InterconnectionabstractTo offer sufficient computing capability, current hardware systems tend to be equipped with multiple chips, while each chip integrates tens to hundreds of cores with Network-on-Chip (NoC) architecture. Thus, chip-to-chip interconnection for NoCs becomes indispensable. However, state-of-the-art inter-chip interconnection, like PCIe or SRIO, suffers from latency and bandwidth bottlenecks especially for communication-intensive tasks. In this paper, we propose a NoC-aware chip-to-chip interconnection scheme. In addition to a lightweight interconnection architecture and protocol, we utilize NoC routers to improve inter-chip connection. A virtual-channel router with transmission priority in the NoC is proposed to eliminate the congestion in the chip-to-chip interconnection. The interconnection system is implemented in Verilog RTL and verified on the Xilinx ZCU102 evaluation kit, achieving up to 10Gb/s/lane line rate with low resource utilization. Wenkang Liao, Yuhao Guo, Shanlin Xiao, Zhiyi Yu |
ISCAS | 3 |
| 2020 | NeuronLink: An Efficient Chip-to-Chip Interconnect for Large-Scale Neural Network AcceleratorsabstractLarge-scale neural network (NN) accelerators typically consist of several processing nodes, which could be implemented as a multi- or many-core chip and organized via a network-on-chip (NoC) to handle the heavy neuron-to-neuron traffic. Multiple NoC-based NN chips are connected through chip-to-chip interconnection networks to further boost the overall neural acceleration capability. Huge amounts of multicast-based traffic travel on-chip or cross chips, making the interconnection network design more challenging and become the bottleneck of the NN system performance and energy. In this article, we propose coupling intrachip and interchip communication techniques, called NeuronLink, for NN accelerators. Regarding the intrachip communication, we propose scoring crossbar arbitration, arbitration interception, and route computation parallelization techniques for virtual-channel routing, leading to a high-throughput NoC with a lower hardware cost for multicast-based traffic. Regarding the interchip communication, we propose a lightweight and NoC-aware chip-to-chip interconnection scheme, enabling efficient interconnection for NoC-based NN chips. In addition, we evaluate the proposed techniques on a four connected NoC-based deep neural network (DNN) chips with four field-programmable gate arrays (FPGAs). The experimental results show that the proposed interconnection network can efficiently manage the data traffic inside DNNs with high-throughput and low-overhead against state-of-the-art interconnects. Shanlin Xiao, Yuhao Guo, Wenkang Liao, Huipeng Deng, Huanliang Zheng, Jian Wang 0080, Gezi Li, Zhiyi Yu |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2016 | An efficient embedded processor for object detection using ASIP methodologyabstractThis paper presents an Application Specific Instruction set Processor (ASIP) for object detection using AdaBoost-based learning algorithm with Haar-like features as weak classifiers. In the proposed ASIP, Single Instruction Multiple Data (SIMD) architecture is adopted for fully exploiting data-level parallelism inherent to the target algorithm.With adding pipeline stages, application-specific registers and custom instructions, AdaBoost algorithm is accelerated by a factor of 13.7× compared to baseline processor. Furthermore, the results show an advantage of the proposed architecture in terms of chip area efficiency while maintain a reliable detection accuracy and achieve real-time object detection at 32fps on VGA video. Shanlin Xiao, Tsuyoshi Isshiki, Dongju Li, Hiroaki Kunieda |
ASAP | 1 |