Dongcheng Zhao

dblp:177/8581 · DBLP profile ↗
← Back
46ranked-venue papers
7as first author
37since 2021 · last 2026
0000-0002-0593-8650ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 25 · 3 first-author · 22 since 2021Systems, architecture and hardware · 12 · 1 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 11 · 9 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 2 since 2021Software engineering, systems software and programming languages · 2 · 1 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author · 2 since 2021Computer networks · 1 · 1 first-author · 1 since 2021Security and privacy · 1 · 1 first-authorHuman-computer interaction and ubiquitous computing · 1
YearPublicationVenuePosition
2026 Boosting the Robustness-Accuracy Trade-off of SNNs by Robust Temporal Self-Ensemble
abstract
Spiking Neural Networks (SNNs) offer a promising direction for energy-efficient and brain-inspired computing, yet their vulnerability to adversarial perturbations remains poorly understood. In this work, we revisit the adversarial robustness of SNNs through the lens of temporal ensembling, treating the network as a collection of evolving sub-networks across discrete timesteps. This formulation uncovers two critical but underexplored challenges—the fragility of individual temporal sub-networks and the tendency for adversarial vulnerabilities to transfer across time. To overcome these limitations, we propose Robust Temporal self-Ensemble (RTE), a training framework that improves the robustness of each sub-network while reducing the temporal transferability of adversarial perturbations. RTE integrates both objectives into a unified loss and employs a stochastic sampling strategy for efficient optimization. Extensive experiments across multiple benchmarks demonstrate that RTE consistently outperforms existing training methods in robust-accuracy trade-off. Additional analyses reveal that RTE reshapes the internal robustness landscape of SNNs, leading to more resilient and temporally diversified decision boundaries. Our study highlights the importance of temporal structure in adversarial learning and offers a principled foundation for building robust spiking models.
Jihang Wang, Dongcheng Zhao, Ruolin Chen, Qian Zhang 0080, Yi Zeng 0001
AAAI2
2026 Hummingbird+: Advancing FPGA-based LLM Deployment from Research Prototype to Edge Product
abstract
Field-Programmable Gate Arrays (FPGAs) have been shown to be viable for Large Language Model (LLM) deployment, but they remain less competitive than embedded GPUs and NPUs for final edge products. This is largely because existing FPGA-based LLM accelerator prototypes rely on large, expensive FPGA devices to provide sufficient hardware resources for satisfactory performance, whereas edge products are highly cost-sensitive. In this work, we move beyond pure architectural prototyping to evaluate the feasibility of using low-cost FPGAs as the final implementation medium for LLM deployment. We propose Hummingbird+, which encompasses: (1) a compact embedded FPGA-based LLM accelerator designed to deliver comparable inference performance compared to embedded GPUs and NPUs, and (2) a custom Printed Circuit Board (PCB) built around a Zynq UltraScale XCZU2CG/3EG SoC, equipped with 24GB of memory and an expected Bill of Materials (BOMs) under \150 in mass production. Through extensive FPGA-centric optimizations, we significantly reduce the accelerator's resource consumption, enabling deployment on entry-level FPGAs with exceptional cost efficiency. On this platform, we successfully deploy the GPTQ 4-bit Qwen3-30B-A3B LLM, achieving a decoding speed of over 18 token/s and a prefill speed of over 50 token/s without further model compression. To our knowledge, this is the first demonstration of an FPGA-based edge product serving as a practical and cost-effective final implementation medium for LLM deployment.
Jindong Li 0001, Guobin Shen, Dongcheng Zhao, Qian Zhang 0080, Yi Zeng 0001
FPGA4
2026 FireFly-P: FPGA-Accelerated Spiking Neural Network Plasticity for Robust Adaptive Control
Jindong Li 0001, Guobin Shen, Dongcheng Zhao, Qian Zhang 0080, Yi Zeng 0001
ISCAS4
2026 Time-efficient task offloading and decentralized collaborative scheduling method for cross-domain computing power networks
Dongcheng Zhao, Guangxia Xu, Gang Sun 0001, Mohsen Guizani
Comput. Networks1
2026 Biologically inspired spiking diffusion model with adaptive lateral selection mechanism
Linghao Feng, Dongcheng Zhao, Sicheng Shen, Yi Zeng 0001
Neural Networks2
2026 FireFly-T: High-Throughput Sparsity Exploitation for Spiking Transformer Acceleration With Dual-Engine Overlay Architecture
abstract
Spiking transformers are emerging as a promising architecture that combines the energy efficiency of Spiking Neural Networks (SNNs) with the powerful attention mechanisms of transformers. However, existing hardware accelerators lack support for spiking attention, exhibit limited throughput when exploiting fine-grained sparsity, and struggle with scalable parallelism in sparse computation. To address these challenges, we propose FireFly-T, a dual-engine overlay architecture that integrates a sparse engine for activation sparsity and a binary engine for spiking attention. In the sparse engine, we present a high-throughput sparse decoder that exploits fine-grained sparsity by concurrently extracting multiple non-zero spikes. To complement this, we introduce a scalable load balancing mechanism with weight dispatch and out-of-order execution, eliminating bank conflicts to support scalable multidimensional parallelism. In the binary engine, we leverage the byte-level write capability of SRAMs to efficiently manipulate the 3D dataflows required for spiking attention with minimal resource overhead. We also optimize the core AND-PopCount operation in spiking attention through a LUT6-based implementation, improving timing closure and reducing LUT utilization on Xilinx FPGAs. As an overlay architecture, FireFly-T further incorporates an orchestrator that dynamically manipulates input dataflows with flexible adaptation to diverse network topologies, while ensuring efficient resource utilization and maintaining high throughput. Experimental results demonstrate that our accelerator achieves 1.39× and 2.40× higher energy efficiency, as well as 4.21× and 7.10× greater DSP efficiency, compared to FireFly v2 and the transformer-enabled SpikeTA, respectively. These results highlight its potential as an efficient hardware platform for spiking transformers.
Jindong Li 0001, Guobin Shen, Dongcheng Zhao, Qian Zhang 0080, Yi Zeng 0001
IEEE Trans. Computers4
2025 EventZoom: A Progressive Approach to Event-Based Data Augmentation for Enhanced Neuromorphic Vision
abstract
Dynamic Vision Sensors (DVS) capture event data with high temporal resolution and low power consumption, presenting a more efficient solution for visual processing in dynamic and real-time scenarios compared to conventional video capture methods. Event data augmentation serves as an essential method for overcoming the limitation of scale and diversity in event datasets. Our comparative experiments demonstrate that the two factors, spatial integrity and temporal continuity, can significantly affect the capacity of event data augmentation, which guarantee the maintenance of the sparsity and high dynamic range characteristics unique to event data. However, existing augmentation methods often neglect the preservation of spatial integrity and temporal continuity. To address this, we developed a novel event data augmentation strategy EventZoom, which employs a temporal progressive strategy, embedding transformed samples into the original samples through progressive scaling and shifting. The scaling process avoids the spatial information loss associated with cropping, while the progressive strategy prevents interruptions or abrupt changes in temporal information. We validated EventZoom across various supervised learning frameworks. The experimental results show that EventZoom consistently outperforms existing event data augmentation methods with SOTA performance. For the first time, we have concurrently employed Semi-supervised and Unsupervised learning to verify feasibility on event augmentation algorithms, demonstrating the applicability and effectiveness of EventZoom as a powerful event-based data augmentation tool in handling real-world scenes with high dynamics and variability environments.
Yiting Dong, Xiang He 0004, Guobin Shen, Dongcheng Zhao, Yang Li 0141, Yi Zeng 0001
AAAI4
2025 StressPrompt: Does Stress Impact Large Language Models and Human Performance Similarly?
abstract
Human beings often experience stress, which can significantly influence their performance. This study explores whether Large Language Models (LLMs) exhibit stress responses similar to those of humans and whether their performance fluctuates under different stress-inducing prompts. To investigate this, we developed a novel set of prompts, termed StressPrompt, designed to induce varying levels of stress. These prompts were derived from established psychological frameworks and carefully calibrated based on ratings from human participants. We then applied these prompts to several LLMs to assess their responses across a range of tasks, including instruction-following, complex reasoning, and emotional intelligence. The findings suggest that LLMs, like humans, perform optimally under moderate stress, consistent with the Yerkes-Dodson law. Notably, their performance declines under both low and high-stress conditions. Our analysis further revealed that these StressPrompts significantly alter the internal states of LLMs, leading to changes in their neural representations that mirror human responses to stress. This research provides critical insights into the operational robustness and flexibility of LLMs, demonstrating the importance of designing AI systems capable of maintaining high performance in real-world scenarios where stress is prevalent, such as in customer service, healthcare, and emergency response contexts. Moreover, this study contributes to the broader AI research community by offering a new perspective on how LLMs handle different scenarios and their similarities to human cognition.
Guobin Shen, Dongcheng Zhao, Aorigele Bao, Xiang He 0004, Yiting Dong, Yi Zeng 0001
AAAI2
2025 Pushing up to the Limit of Memory Bandwidth and Capacity Utilization for Efficient LLM Decoding on Embedded FPGA
abstract
The extremely high computational and storage demands of large language models have excluded most edge devices, which were widely used for efficient machine learning, from being viable options. A typical edge device usually only has 4GB of memory capacity and a bandwidth of less than 20GB/s, while a large language model quantized to 4-bit precision with 7B parameters already requires 3.5GB of capacity, and its decoding process is purely bandwidth-bound. In this paper, we aim to explore these limits by proposing a hardware accelerator for large language model (LLM) inference on the Zynq-based KV260 platform, equipped with 4GB of 64-bit 2400Mbps DDR4 memory. We successfully deploy a LLaMA2-7B model, achieving a decoding speed of around 5 token/s, utilizing 93.3% of the memory capacity and reaching 85% decoding speed of the theoretical memory bandwidth limit. To fully reserve the memory capacity for model weights and key-value cache, we develop the system in a bare-metal environment without an operating system. To fully reserve the bandwidth for model weight transfers, we implement a customized dataflow with an operator fusion pipeline and propose a data arrangement format that can maximize the data transaction efficiency. This research marks the first attempt to deploy a 7B level LLM on a standalone embedded field programmable gate array (FPGA) device. It provides key insights into efficient LLM inference on embedded FPGA devices and provides guidelines for future architecture design.
Jindong Li 0001, Guobin Shen, Dongcheng Zhao, Qian Zhang 0080, Yi Zeng 0001
DATE4
2025 Hummingbird: A Smaller and Faster Large Language Model Accelerator on Embedded FPGA
abstract
Deploying large language models (LLMs) on embedded devices remains a significant research challenge due to the high computational and memory demands of LLMs and the limited hardware resources available in such environments. While embedded FPGAs have demonstrated performance and energy efficiency in traditional deep neural networks, their potential for LLM inference remains largely unexplored. Recent efforts to deploy LLMs on FPGAs have primarily relied on large, expensive cloud-grade hardware and have only shown promising results on relatively small LLMs, limiting their real-world applicability. In this work, we present Hummingbird, a novel FPGA accelerator designed specifically for LLM inference on embedded FPGAs. Hummingbird is smaller—targeting embedded FPGAs such as the KV260 and ZCU104 with 67% LUT, 39% DSP, and 42% power savings over existing research. Hummingbird is stronger—targeting LLaMA3-8B and supporting longer contexts, overcoming the typical 4GB memory constraint of embedded FPGAs through offloading strategies. Finally, Hummingbird is faster—achieving 4.8 tokens/s and 8.6 tokens/s for LLaMA3-8B on the KV260 and ZCU104 respectively, with 93-94% model bandwidth utilization, outperforming the prior 4.9 token/s for LLaMA2-7B with 84% bandwidth utilization baseline. We further demonstrate the viability of industrial applications by deploying Hummingbird on a cost-optimized Spartan UltraScale FPGA, paving the way for affordable LLM solutions at the edge.
Jindong Li 0001, Ruiqi Chen 0001, Guobin Shen, Dongcheng Zhao, Qian Zhang 0080, Yi Zeng 0001
ICCAD5
2025 SpikePack: Enhanced Information Flow in Spiking Neural Networks with High Hardware Compatibility
abstract
Spiking Neural Networks (SNNs) hold promise for energy-efficient, biologically inspired computing. We identify substantial informatio loss during spike transmission, linked to temporal dependencies in traditional Leaky Integrate-and-Fire (LIF) neuron-a key factor potentially limiting SNN performance. Existing SNN architectures also underutilize modern GPUs, constrained by single-bit spike storage and isolated weight-spike operations that restrict computational efficiency. We introduce ${SpikePack}$, a neuron model designed to reduce transmission loss while preserving essential features like membrane potential reset and leaky integration. ${SpikePack}$ achieves constant $\mathcal{O}(1)$ time and space complexity, enabling efficient parallel processing on GPUs and also supporting serial inference on existing SNN hardware accelerators. Compatible with standard Artificial Neural Network (ANN) architectures, ${SpikePack}$ facilitates near-lossless ANN-to-SNN conversion across various networks. Experimental results on tasks such as image classification, detection, and segmentation show ${SpikePack}$ achieves significant gains in accuracy and efficiency for both directly trained and converted SNNs over state-of-the-art models. Tests on FPGA-based platforms further confirm cross-platform flexibility, delivering high performance and enhanced sparsity. By enhancing information flow and rethinking SNN-ANN integration, ${SpikePack}$ advances efficient SNN deployment across diverse hardware platforms.
Guobin Shen, Jindong Li 0001, Dongcheng Zhao, Yi Zeng 0001
ICCV4
2025 Jailbreak Antidote: Runtime Safety-Utility Balance via Sparse Representation Adjustment in Large Language Models
abstract
As large language models (LLMs) become integral to various applications, ensuring both their safety and utility is paramount. Jailbreak attacks, which manipulate LLMs into generating harmful content, pose significant challenges to this balance. Existing defenses, such as prompt engineering and safety fine-tuning, often introduce computational overhead, increase inference latency, and lack runtime flexibility. Moreover, overly restrictive safety measures can degrade model utility by causing refusals of benign queries. In this paper, we introduce *Jailbreak Antidote*, a method that enables real-time adjustment of LLM safety preferences by manipulating a sparse subset of the model's internal states during inference. By shifting the model's hidden representations along a safety direction with varying strengths, we achieve flexible control over the safety-utility balance without additional token overhead or inference delays. Our analysis reveals that safety-related information in LLMs is sparsely distributed; adjusting approximately *5\%* of the internal state is as effective as modifying the entire state. Extensive experiments on nine LLMs (ranging from 2 billion to 72 billion parameters), evaluated against ten jailbreak attack methods and compared with six defense strategies, validate the effectiveness and efficiency of our approach. By directly manipulating internal states during reasoning, *Jailbreak Antidote* offers a lightweight, scalable solution that enhances LLM safety while preserving utility, opening new possibilities for real-time safety mechanisms in widely-deployed AI systems.
Guobin Shen, Dongcheng Zhao, Yiting Dong, Xiang He 0004, Yi Zeng 0001
ICLR2
2025 Brain-Inspired Stepwise Patch Merging for Vision Transformers
abstract
The hierarchical architecture has become a mainstream design paradigm for Vision Transformers (ViTs), with Patch Merging serving as the pivotal component that transforms a columnar architecture into a hierarchical one. Drawing inspiration from the brain's ability to integrate global and local information for comprehensive visual understanding, we propose Stepwise Patch Merging (SPM), which enhances the subsequent attention mechanism's ability to 'see' better. SPM consists of Multi-Scale Aggregation (MSA) and Guided Local Enhancement (GLE) striking a proper balance between long-range dependency modeling and local feature enhancement. Extensive experiments conducted on benchmark datasets, including ImageNet-1K, COCO, and ADE20K, demonstrate that SPM significantly improves the performance of various models, particularly in dense prediction tasks such as object detection and semantic segmentation. Meanwhile, experiments show that combining SPM with different backbones can further improve performance. The code has been released at https://github.com/Yonghao-Yu/StepwisePatchMerging.
Dongcheng Zhao, Guobin Shen, Yiting Dong, Yi Zeng 0001
IJCAI2
2025 Learning the Plasticity: Plasticity-Driven Learning Framework in Spiking Neural Networks
abstract
The evolution of the human brain has led to the development of complex synaptic plasticity, enabling dynamic adaptation to a constantly evolving world. This progress inspires our exploration into a new paradigm for Spiking Neural Networks (SNNs): a Plasticity-Driven Learning Framework (PDLF). This paradigm diverges from traditional neural network models that primarily focus on direct training of synaptic weights, leading to static connections that limit adaptability in dynamic environments. Instead, our approach delves into the heart of synaptic behavior, prioritizing the learning of plasticity rules themselves. This shift in focus from weight adjustment to mastering the intricacies of synaptic change offers a more flexible and dynamic pathway for neural networks to evolve and adapt. Our PDLF does not merely adapt existing concepts of functional and Presynaptic-Dependent Plasticity but redefines them, aligning closely with the dynamic and adaptive nature of biological learning. This reorientation enhances key cognitive abilities in artificial intelligence systems, such as working memory and multitasking capabilities, and demonstrates superior adaptability in complex, real-world scenarios. Moreover, our framework sheds light on the intricate relationships between various forms of plasticity and cognitive functions, thereby contributing to a deeper understanding of the brain's learning mechanisms. Integrating this groundbreaking plasticity-centric approach in SNNs marks a significant advancement in the fusion of neuroscience and artificial intelligence. It paves the way for developing AI systems that not only learn but also adapt in an ever-changing world, much like the human brain.
Guobin Shen, Dongcheng Zhao, Yiting Dong, Yang Li 0141, Yi Zeng 0001
NeurIPS2
2025 STEP: A Unified Spiking Transformer Evaluation Platform for Fair and Reproducible Benchmarking
abstract
Spiking Transformers have recently emerged as promising architectures for combining the efficiency of spiking neural networks with the representational power of self-attention. However, the lack of standardized implementations, evaluation pipelines, and consistent design choices has hindered fair comparison and principled analysis. In this paper, we introduce \textbf{STEP}, a unified benchmark framework for Spiking Transformers that supports a wide range of tasks, including classification, segmentation, and detection across static, event-based, and sequential datasets. STEP provides modular support for diverse components such as spiking neurons, input encodings, surrogate gradients, and multiple backends (e.g., SpikingJelly, BrainCog). Using STEP, we reproduce and evaluate several representative models, and conduct systematic ablation studies on attention design, neuron types, encoding schemes, and temporal modeling capabilities. We also propose a unified analytical model for energy estimation, accounting for spike sparsity, bitwidth, and memory access, and show that quantized ANNs may offer comparable or better energy efficiency. Our results suggest that current Spiking Transformers rely heavily on convolutional frontends and lack strong temporal modeling, underscoring the need for spike-native architectural innovations. The full code is available at: https://github.com/Fancyssc/STEP.
Sicheng Shen, Dongcheng Zhao, Linghao Feng, Zeyang Yue, Jindong Li 0001, Guobin Shen, Yi Zeng 0001
NeurIPS2
2025 Improving stability and performance of spiking neural networks through enhancing temporal consistency
Dongcheng Zhao, Guobin Shen, Yiting Dong, Yang Li 0141, Yi Zeng 0001
Pattern Recognit.1
2025 FireFly-S: Exploiting Dual-Side Sparsity for Spiking Neural Networks Acceleration With Reconfigurable Spatial Architecture
abstract
Spiking Neural Networks (SNNs), with their brain-inspired structure using discrete spikes instead of continuous activations, are gaining attention for their potential of efficient processing on neuromorphic chips. While current SNN hardware accelerators often prioritize temporal spike sparsity, exploiting sparse synaptic weights offers significant untapped potential for even greater efficiency. To address this, we propose FireFly-S, a Sparse extension of the FireFly series. This co-optimized software-hardware design focusing on leveraging dual-side sparsity for acceleration. On the software side, we propose a novel algorithmic optimization framework that combines gradient rewiring for pruning and modified Learned Step Size Quantization (LSQ) tailored for SNNs, which achieves remarkable weight sparsity exceeding 85% and enables efficient 4-bit quantization with negligible accuracy loss. On the hardware side, we present an efficient dual-side sparsity detector employing a Bitmap-based sparse decoding logic to pinpoint the positions of non-zero weights and input spikes. The logic allows for the direct bypassing of redundant computations, thereby enhancing computational efficiency. Different from the overlay architecture adopted by previous FireFly series, we adopt a parametric spatial architecture with inter-layer pipelining that can fully exploit the fine-grained programmability and reconfigurability of Field-Programmable Gate Arrays (FPGAs), enabling fast deployment for various models. A spatial-temporal dataflow is also proposed to support such inter-layer pipelining and avoid long-term temporal dependencies. In experiments conducted on the MNIST, DVS-Gesture and CIFAR-10 datasets, the FireFly-S model achieves 85-95% sparsity with 4-bit quantization and the hardware accelerator effectively leverages the dual-side sparsity, delivering outstanding performance metrics of 10,047 FPS/W on MNIST, 3,683 FPS/W on DVS-Gesture, and 2,327 FPS/W on CIFAR-10.
Jindong Li 0001, Guobin Shen, Dongcheng Zhao, Qian Zhang 0080, Yi Zeng 0001
IEEE Trans. Circuits Syst. I Regul. Pap.4
2024 An Efficient Knowledge Transfer Strategy for Spiking Neural Networks from Static to Event Domain
abstract
Spiking neural networks (SNNs) are rich in spatio-temporal dynamics and are suitable for processing event-based neuromorphic data. However, event-based datasets are usually less annotated than static datasets. This small data scale makes SNNs prone to overfitting and limits their performance. In order to improve the generalization ability of SNNs on event-based datasets, we use static images to assist SNN training on event data. In this paper, we first discuss the domain mismatch problem encountered when directly transferring networks trained on static datasets to event data. We argue that the inconsistency of feature distributions becomes a major factor hindering the effective transfer of knowledge from static images to event data. To address this problem, we propose solutions in terms of two aspects: feature distribution and training strategy. Firstly, we propose a knowledge transfer loss, which consists of domain alignment loss and spatio-temporal regularization. The domain alignment loss learns domain-invariant spatial features by reducing the marginal distribution distance between the static image and the event data. Spatio-temporal regularization provides dynamically learnable coefficients for domain alignment loss by using the output features of the event data at each time step as a regularization term. In addition, we propose a sliding training strategy, which gradually replaces static image inputs probabilistically with event data, resulting in a smoother and more stable training for the network. We validate our method on neuromorphic datasets, including N-Caltech101, CEP-DVS, and N-Omniglot. The experimental results show that our proposed method achieves better performance on all datasets compared to the current state-of-the-art methods. Code is available at https://github.com/Brain-Cog-Lab/Transfer-for-DVS.
Xiang He 0004, Dongcheng Zhao, Yang Li 0141, Guobin Shen, Qingqun Kong, Yi Zeng 0001
AAAI2
2024 Are Conventional SNNs Really Efficient? A Perspective from Network Quantization
abstract
Spiking Neural Networks (SNNs) have been widely praised for their high energy efficiency and immense potential. However, comprehensive research that critically contrasts and correlates SNNs with quantized Artificial Neural Networks (ANNs) remains scant, often leading to skewed comparisons lacking fairness towards ANNs. This paper introduces a unified perspective, illustrating that the time steps in SNNs and quantized bit-widths of activation values present analogous representations. Building on this, we present a more pragmatic and rational approach to estimating the energy consumption of SNNs. Diverging from the conventional Synaptic Operations (SynOps), we champion the “Bit Budget” concept. This notion permits an intricate discourse on strategically allocating computational and storage resources between weights, activation values, and temporal steps under stringent hardware constraints. Guided by the Bit Budget paradigm, we discern that pivoting efforts towards spike patterns and weight quantization, rather than temporal attributes, elicits profound implications for model performance. Utilizing the Bit Budget for holistic design consideration of SNNs elevates model performance across diverse data types, encompassing static imagery and neuromorphic datasets. Our revelations bridge the theoretical chasm between SNNs and quantized ANNs and illuminate a pragmatic trajectory for future endeavors in energy-efficient neural computations.
Guobin Shen, Dongcheng Zhao, Jindong Li 0001, Yi Zeng 0001
CVPR2
2024 Revealing Untapped DSP Optimization Potentials for FPGA-Based Systolic Matrix Engines
abstract
Systolic architectures are widely embraced by neural network accelerators for their superior performance in highly parallelized computation. The DSP48E2s serve as dedicated arithmetic blocks in Xilinx Ultrascale series FPGAs and constitute a fundamental component in FPGA-based systolic matrix engines. Harnessing the full potential of DSP48E2s in architectural design can result in significant performance enhancements for systolic architectures on Ultrascale series FPGAs. This paper unveils several previously untapped DSP optimization techniques capable of further enhancing FPGA-based systolic matrix engines. We apply these techniques to two well-known systolic architectures: Google TPUv1 and Xilinx Vitis AI DPU. With the proposed techniques, our design achieves substantial resource and power reduction compared to the open-source TPUv1 FPGA implementation and the Vitis AI DPU implementation in the same parallelism setting. We also demonstrate the applicability of our techniques to neuromorphic hardware for supporting spiking neural network acceleration.
Jindong Li 0001, Guobin Shen, Dongcheng Zhao, Qian Zhang 0080, Yi Zeng 0001
FPL4
2024 TIM: An Efficient Temporal Interaction Module for Spiking Transformer
Sicheng Shen, Dongcheng Zhao, Guobin Shen, Yi Zeng 0001
IJCAI2
2024 Parallel Spiking Unit for Efficient Training of Spiking Neural Networks
abstract
Efficient parallel computing has become a pivotal element in advancing artificial intelligence. Yet, the deployment of Spiking Neural Networks (SNNs) in this domain is hampered by their inherent sequential computational dependency. This constraint arises from the need for each time step’s processing to rely on the preceding step’s outcomes, significantly impeding the adaptability of SNN models to massively parallel computing environments. Addressing this challenge, our paper introduces the innovative Parallel Spiking Unit (PSU) and its two derivatives, the Input-aware PSU (IPSU) and Reset-aware PSU (RPSU). These variants skillfully decouple the leaky integration and firing mechanisms in spiking neurons while probabilistically managing the reset process. By preserving the fundamental computational attributes of the spiking neuron model, our approach enables the concurrent computation of all membrane potential instances within the SNN, facilitating parallel spike output generation and substantially enhancing computational efficiency. Comprehensive testing across various datasets, including static and sequential images, Dynamic Vision Sensor (DVS) data, and speech datasets, demonstrates that the PSU and its variants not only significantly boost performance and simulation speed but also augment the energy efficiency of SNNs through enhanced sparsity in neural activity. These advancements underscore the potential of our method in revolutionizing SNN deployment for high-performance parallel computing applications.
Yang Li 0141, Yinqian Sun, Xiang He 0004, Yiting Dong, Dongcheng Zhao, Yi Zeng 0001
IJCNN5
2024 CACE-Net: Co-guidance Attention and Contrastive Enhancement for Effective Audio-Visual Event Localization
Xiang He 0004, Xiangxi Liu, Yang Li 0141, Dongcheng Zhao, Guobin Shen, Qingqun Kong, Xin Yang 0001, Yi Zeng 0001
ACM Multimedia4
2024 Neuro-Vision to Language: Enhancing Brain Recording-based Visual Reconstruction and Language Interaction
abstract
Decoding non-invasive brain recordings is pivotal for advancing our understanding of human cognition but faces challenges due to individual differences and complex neural signal representations. Traditional methods often require customized models and extensive trials, lacking interpretability in visual reconstruction tasks. Our framework integrates 3D brain structures with visual semantics using a *Vision Transformer 3D*. This unified feature extractor efficiently aligns fMRI features with multiple levels of visual embeddings, eliminating the need for subject-specific models and allowing extraction from single-trial data. The extractor consolidates multi-level visual features into one network, simplifying integration with Large Language Models (LLMs). Additionally, we have enhanced the fMRI dataset with diverse fMRI-image-related textual data to support multimodal large model development. Integrating with LLMs enhances decoding capabilities, enabling tasks such as brain captioning, complex reasoning, concept localization, and visual reconstruction. Our approach demonstrates superior performance across these tasks, precisely identifying language-based concepts within brain signals, enhancing interpretability, and providing deeper insights into neural processes. These advances significantly broaden the applicability of non-invasive brain decoding in neuroscience and human-computer interaction, setting the stage for advanced brain-computer interfaces and cognitive models.
Guobin Shen, Dongcheng Zhao, Xiang He 0004, Linghao Feng, Yiting Dong, Jihang Wang, Qian Zhang 0080, Yi Zeng 0001
NeurIPS2
2024 MSAT: biologically inspired multistage adaptive threshold for conversion of spiking neural networks
Xiang He 0004, Yang Li 0141, Dongcheng Zhao, Qingqun Kong, Yi Zeng 0001
Neural Comput. Appl.3
2024 Spiking generative adversarial network with attention scoring decoding
Linghao Feng, Dongcheng Zhao, Yi Zeng 0001
Neural Networks2
2024 Directly training temporal Spiking Neural Network with sparse surrogate gradient
Yang Li 0141, Dongcheng Zhao, Yi Zeng 0001
Neural Networks3
2024 Exploiting nonlinear dendritic adaptive computation in training deep Spiking Neural Networks
abstract
Inspired by the information transmission process in the brain, Spiking Neural Networks (SNNs) have gained considerable attention due to their event-driven nature. However, as the network structure grows complex, managing the spiking behavior within the network becomes challenging. Networks with excessively dense or sparse spikes fail to transmit sufficient information, inhibiting SNNs from exhibiting superior performance. Current SNNs linearly sum presynaptic information in postsynaptic neurons, overlooking the adaptive adjustment effect of dendrites on information processing. In this study, we introduce the Dendritic Spatial Gating Module (DSGM), which scales and translates the input, reducing the loss incurred when transforming the continuous membrane potential into discrete spikes. Simultaneously, by implementing the Dendritic Temporal Adjust Module (DTAM), dendrites assign different importance to inputs of different time steps, facilitating the establishment of the temporal dependency of spiking neurons and effectively integrating multi-step time information. The fusion of these two modules results in a more balanced spike representation within the network, significantly enhancing the neural network's performance. This approach has achieved state-of-the-art performance on static image datasets, including CIFAR10 and CIFAR100, as well as event datasets like DVS-CIFAR10, DVS-Gesture, and N-Caltech101. It also demonstrates competitive performance compared to the current state-of-the-art on the ImageNet dataset.
Guobin Shen, Dongcheng Zhao, Yi Zeng 0001
Neural Networks2
2024 FireFly v2: Advancing Hardware Support for High-Performance Spiking Neural Network With a Spatiotemporal FPGA Accelerator
abstract
Spiking Neural Networks (SNNs) are expected to be a promising alternative to Artificial Neural Networks (ANNs) due to their strong biological interpretability and high energy efficiency. Specialized SNN hardware offers clear advantages over general-purpose devices in terms of power and performance. However, there’s still room to advance hardware support for state-of-the-art (SOTA) SNN algorithms and improve computation and memory efficiency. As a further step in supporting high-performance SNNs on specialized hardware, we introduce FireFly v2, an FPGA SNN accelerator that can address the issue of non-spike operation in current SOTA SNN algorithms, which presents an obstacle in the end-to-end deployment onto existing SNN hardware. To more effectively align with the SNN characteristics, we design a spatiotemporal dataflow that allows four dimensions of parallelism and eliminates the need for membrane potential storage, enabling on-the-fly spike processing and spike generation. To further improve hardware acceleration performance, we develop a high-performance spike computing engine as a backend based on a systolic array operating at 500-600MHz. To the best of our knowledge, FireFly v2 achieves the highest clock frequency among all FPGA-based implementations. Furthermore, it stands as the first SNN accelerator capable of supporting non-spike operations, which are commonly used in advanced SNN algorithms. FireFly v2 has doubled the throughput and DSP efficiency when compared to our previous version of FireFly and it exhibits ×1.33 the DSP efficiency and ×1.42 the power efficiency compared to the current most advanced FPGA accelerators.
Jindong Li 0001, Guobin Shen, Dongcheng Zhao, Qian Zhang 0080, Yi Zeng 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3
2023 Bullying10K: A Large-Scale Neuromorphic Dataset towards Privacy-Preserving Bullying Recognition
abstract
The prevalence of violence in daily life poses significant threats to individuals' physical and mental well-being. Using surveillance cameras in public spaces has proven effective in proactively deterring and preventing such incidents. However, concerns regarding privacy invasion have emerged due to their widespread deployment.To address the problem, we leverage Dynamic Vision Sensors (DVS) cameras to detect violent incidents and preserve privacy since it captures pixel brightness variations instead of static imagery. We introduce the Bullying10K dataset, encompassing various actions, complex movements, and occlusions from real-life scenarios. It provides three benchmarks for evaluating different tasks: action recognition, temporal action localization, and pose estimation. With 10,000 event segments, totaling 12 billion events and 255 GB of data, Bullying10K contributes significantly by balancing violence detection and personal privacy persevering. And it also poses a challenge to the neuromorphic dataset. It will serve as a valuable resource for training and developing privacy-protecting video systems. The Bullying10K opens new possibilities for innovative approaches in these domains.
Yiting Dong, Yang Li 0141, Dongcheng Zhao, Guobin Shen, Yi Zeng 0001
NeurIPS3
2023 EventMix: An efficient data augmentation strategy for event-based learning
abstract
High-quality and challenging event stream datasets play an important role in the design of an efficient event-driven mechanism that mimics the brain. Although event cameras can provide high dynamic range and low-energy event stream data, the scale is smaller and more difficult to obtain than traditional frame-based data, which restricts the development of neuromorphic computing. Data augmentation can improve the quantity and quality of the original data by processing more representations from the original data. This paper proposes an efficient data augmentation strategy for event stream data: EventMix. We carefully design the mixing of different event streams by Gaussian Mixture Model (GMM) to generate random 3D masks and achieve arbitrary shape mixing of event streams in the spatio-temporal dimension. By computing the relative distances of event streams, we propose a more reasonable way to assign labels to the mixed samples. The experimental results on multiple neuromorphic datasets have shown that our strategy can improve performance on neuromorphic classification tasks as well as neuromorphic human action recognition tasks both for ANNs and SNNs, and we have achieved state-of-the-art performance on DVS-CIFAR10, N-Caltech101, and DVS-Gesture datasets.
Guobin Shen, Dongcheng Zhao, Yi Zeng 0001
Inf. Sci.2
2023 An unsupervised STDP-based spiking neural network inspired by biologically plausible learning rules and connections
abstract
The backpropagation algorithm has promoted the rapid development of deep learning, but it relies on a large amount of labeled data and still has a large gap with how humans learn. The human brain can quickly learn various conceptual knowledge in a self-organized and unsupervised manner, accomplished through coordinating various learning rules and structures in the human brain. Spike-timing-dependent plasticity (STDP) is a general learning rule in the brain, but spiking neural networks (SNNs) trained with STDP alone is inefficient and perform poorly. In this paper, taking inspiration from short-term synaptic plasticity, we design an adaptive synaptic filter and introduce the adaptive spiking threshold as the neuron plasticity to enrich the representation ability of SNNs. We also introduce an adaptive lateral inhibitory connection to adjust the spikes balance dynamically to help the network learn richer features. To speed up and stabilize the training of unsupervised spiking neural networks, we design a samples temporal batch STDP (STB-STDP), which updates weights based on multiple samples and moments. By integrating the above three adaptive mechanisms and STB-STDP, our model greatly accelerates the training of unsupervised spiking neural networks and improves the performance of unsupervised SNNs on complex tasks. Our model achieves the current state-of-the-art performance of unsupervised STDP-based SNNs in the MNIST and FashionMNIST datasets. Further, we tested on the more complex CIFAR10 dataset, and the results fully illustrate the superiority of our algorithm. Our model is also the first work to apply unsupervised STDP-based SNNs to CIFAR10. At the same time, in the small-sample learning scenario, it will far exceed the supervised ANN using the same structure.
Yiting Dong, Dongcheng Zhao, Yang Li 0141, Yi Zeng 0001
Neural Networks2
2023 FireFly: A High-Throughput Hardware Accelerator for Spiking Neural Networks With Efficient DSP and Memory Optimization
abstract
Spiking neural networks (SNNs) have been widely used due to their strong biological interpretability and high-energy efficiency. With the introduction of the backpropagation algorithm and surrogate gradient, the structure of SNNs has become more complex, and the performance gap with artificial neural networks (ANNs) has gradually decreased. However, most SNN hardware implementations for field-programmable gate arrays (FPGAs) cannot meet arithmetic or memory efficiency requirements, which significantly restricts the development of SNNs. They do not delve into the arithmetic operations between the binary spikes and synaptic weights or assume unlimited on-chip RAM resources using overly expensive devices on small tasks. To improve arithmetic efficiency, we analyze the neural dynamics of spiking neurons, generalize the SNN arithmetic operation to the multiplex-accumulate operation, and propose a high-performance implementation of such operation by utilizing the DSP48E2 hard block in Xilinx Ultrascale FPGAs. To improve memory efficiency, we design a memory system to enable efficient synaptic weights and membrane voltage memory access with reasonable on-chip RAM consumption. Combining the above two improvements, we propose an FPGA accelerator that can process spikes generated by the firing neurons on-the-fly (FireFly). FireFly is the first SNN accelerator that incorporates DSP optimization techniques into SNN synaptic operations. FireFly is implemented on several FPGA edge devices with limited resources but still guarantees a peak performance of 5.53 TOP/s at 300 MHz. As a lightweight accelerator, FireFly achieves the highest computational density efficiency compared with existing research using large FPGA devices.
Jindong Li 0001, Guobin Shen, Dongcheng Zhao, Qian Zhang 0080, Yi Zeng 0001
IEEE Trans. Very Large Scale Integr. Syst.3
2022 The Processing Method of the Message Based on the In-band Network Telemetry Technology
abstract
With the innovation of business applications and the continuous growth of user scale, the network presents the characteristics of "high-speed, large-scale, multi access and unpredictable". The management and control methods and means of the traditional network have been difficult to solve the challenges of existing networks and future networks. Therefore, the network managers urgently need to subvert the monitoring and troubleshooting methods of the traditional network and put forward real-time and flexible measurement solutions that can deal with scenario use cases such as the network state measurement, network failure detection, fault location and recovery. Therefore, the paper proposes a system and method for how the In-band Network Telemetry (INT) technology of switching equipment is applied in complex public packet networks, to deal with scenario use cases such as the network state measurement, network failure detection, fault location and recovery. However, in a cross-domain network, it is difficult to implement the network telemetry in a cross-domain network. Therefore, this paper proposes a method to realize the network telemetry in the cross-domain network (the complex public packet networks) by using the tunneling technology to deal with scenario use cases such as the network state measurement, network failure detection, fault location and recovery.
Congcong Min, Dongcheng Zhao, Hua Lu 0012
ICSS2
2022 A Machine Learning Method and Device Based on Programmable Switch
abstract
Machine learning methods have many excellent properties, such as the quality and efficiency of algorithmic that increase with the number of training sessions as new data is fed in. In a large network topology, how to balance the traffic in the network and improve the link utilization has always been a concern of network engineers. We take advantage of the Programming Protocol-Independent Packet Processors(P4) language with protocol-independent features, use in-band telemetry to collect port traffic statistics, delays and other information on the link, use routing protocols to collect topology, weight etc and transmit these information to machine learning server through the Google Remote Procedure Calls(gRPC) interface. The machine learning server uses a machine learning algorithm to generate a policy for adjusting the traffic, converts the policy into a forwarding table and sends it to the forwarding plane to balance the link traffic.
Congcong Min, Dongcheng Zhao, Hua Lu 0012
ICSS2
2022 Spiking CapsNet: A spiking neural network with a biologically plausible routing rule between capsules
abstract
Spiking neural network (SNN) has attracted much attention due to its powerful spatio-temporal information representation ability. Capsule Neural Network (CapsNet) does well in assembling and coupling features of different network layers. Here, we propose Spiking CapsNet by combining spiking neurons and capsule structures. In addition, we propose a more biologically plausible Spike Timing Dependent Plasticity routing mechanism. The coupling ability is further improved by fully considering the spatio-temporal relationship between spiking capsules of the low layer and the high layer. We have verified experiments on the MNIST, FashionMNIST, and CIFAR10 datasets. Our algorithm still shows comparable performance concerning other excellent SNNs with typical structures (convolutional, fully-connected) on these classification tasks. Our Spiking CapsNet combines SNN and CapsNet’s strengths and shows strong robustness to noise and affine transformation. By adding different Salt-Pepper and Gaussian noise to the test dataset, the experimental results demonstrate that our algorithm is more resistant to noise than other approaches. As well, our Spiking CapsNet shows strong generalization to affine transformation on the AffNIST dataset. Our code is available at https://github.com/BrainCog-X/Brain-Cog.
Dongcheng Zhao, Yang Li 0141, Yi Zeng 0001, Jihang Wang, Qian Zhang 0080
Inf. Sci.1
2022 BackEISNN: A deep spiking neural network with adaptive self-feedback and balanced excitatory-inhibitory neurons
abstract
Spiking neural networks (SNNs) transmit information through discrete spikes that perform well in processing spatial-temporal information. Owing to their nondifferentiable characteristic, difficulties persist in designing SNNs that deliver good performance. SNNs trained with backpropagation have recently exhibited impressive performance by using gradient approximation. However, their performance on complex tasks remains significantly inferior to that of deep neural networks. By taking inspiration from autapses in the brain that connect spiking neurons with a self-feedback connection, we apply adaptive time-delayed self-feedback to the membrane potential to regulate the precision of the spikes. We also strike a balance between the excitatory and inhibitory mechanisms of neurons to dynamically control the output of spiking neurons. By combining these two mechanisms, we propose a deep SNN with adaptive self-feedback and balanced excitatory and inhibitory neurons (BackEISNN). The results of experiments on several standard datasets show that the two modules not only accelerate the convergence of the network but also increase its accuracy. Our model achieved state-of-the-art performance on the MNIST, Fashion-MNIST, and N-MNIST datasets. The proposed BackEISNN also achieved remarkably good performance on the CIFAR10 dataset while using a relatively light structure that competes against state-of-the-art SNNs.
Dongcheng Zhao, Yi Zeng 0001, Yang Li 0141
Neural Networks1
2019 Dynamic Fusion of Convolutional Features based on Spatial and Temporal Attention for Visual Tracking
abstract
Convolutional neural networks (CNN) based trackers have been widely employed in visual object tracking due to their powerful representations. Features from different CNN layers encode different information. Deeper layers contain more semantic information, while the resolution is too coarse to localize the target. Shallower layers carry more detail information but are less robust for appearance variations. In this paper, we propose an algorithm which incorporates the Spatial and Temporal attention to take full advantage of the Hierarchical Convolutional Features for Tracking (STHCFT). We firstly learn correlation filters on each convolutional layer. Based on the spatial attention inspired by the paraventricular thalamus (PVT) in the brain, we choose the most important layer to build the base response, and the others to be the auxiliary responses. In addition, we make full use of the temporal attention to determine the weights of the auxiliary responses. Finally, the target is located by the maximum value of the fused responses. Extensive experimental results on the benchmark OTB-2013 and OTB-2015 have shown the proposed algorithm performs favorably against several state-of-the-art trackers.
Dongcheng Zhao, Yi Zeng 0001
IJCNN1
2019 Mobile-aware service function chain migration in cloud-fog computing
Dongcheng Zhao, Gang Sun 0001, Dan Liao, Shizhong Xu, Victor Chang 0001
Future Gener. Comput. Syst.1
2018 A Plasticity-Centric Approach to Train the Non-Differential Spiking Neural Networks
abstract
Many efforts have been taken to train spiking neural networks (SNNs), but most of them still need improvements due to the discontinuous and non-differential characteristics of SNNs. While the mammalian brains solve these kinds of problems by integrating a series of biological plasticity learning rules. In this paper, we will focus on two biological plausible methodologies and try to solve these catastrophic training problems in SNNs. Firstly, the biological neural network will try to keep a balance between inputs and outputs on both the neuron and the network levels. Secondly, the biological synaptic weights will be passively updated by the changes of the membrane potentials of the neighbour-hood neurons, and the plasticity of synapses will not propagate back to other previous layers. With these biological inspirations, we propose Voltage-driven Plasticity-centric SNN (VPSNN), which includes four steps, namely: feed forward inference, unsupervised equilibrium state learning, supervised last layer learning and passively updating synaptic weights based on spike-timing dependent plasticity (STDP). Finally we get the accuracy of 98.52% on the hand-written digits classification task on MNIST. In addition, with the help of a visualization tool, we try to analyze the black box of SNN and get better understanding of what benefits have been acquired by the proposed method.
Tielin Zhang, Yi Zeng 0001, Dongcheng Zhao, Mengting Shi
AAAI3
2018 Brain-inspired Balanced Tuning for Spiking Neural Networks
abstract
Due to the nature of Spiking Neural Networks (SNNs), it is challenging to be trained by biologically plausible learning principles. The multi-layered SNNs are with non-differential neurons, temporary-centric synapses, which make them nearly impossible to be directly tuned by back propagation. Here we propose an alternative biological inspired balanced tuning approach to train SNNs. The approach contains three main inspirations from the brain: Firstly, the biological network will usually be trained towards the state where the temporal update of variables are equilibrium (e.g. membrane potential); Secondly, specific proportions of excitatory and inhibitory neurons usually contribute to stable representations; Thirdly, the short-term plasticity (STP) is a general principle to keep the input and output of synapses balanced towards a better learning convergence. With these inspirations, we train SNNs with three steps: Firstly, the SNN model is trained with three brain-inspired principles; then weakly supervised learning is used to tune the membrane potential in the final layer for network classification; finally the learned information is consolidated from membrane potential into the weights of synapses by Spike-Timing Dependent Plasticity (STDP). The proposed approach is verified on the MNIST hand-written digit recognition dataset and the performance (the accuracy of 98.64%) indicates that the ideas of balancing state could indeed improve the learning ability of SNNs, which shows the power of proposed brain-inspired approach on the tuning of biological plausible SNNs.
Tielin Zhang, Yi Zeng 0001, Dongcheng Zhao, Bo Xu 0002
IJCAI3
2018 Towards provisioning hybrid virtual networks in federated cloud data centers
Gang Sun 0001, Dan Liao, Dongcheng Zhao, Zhili Sun, Victor Chang 0001
Future Gener. Comput. Syst.3
2018 Live Migration for Multiple Correlated Virtual Machines in Cloud-Based Data Centers
abstract
With the development of cloud computing, virtual machine migration is emerging as a promising technique to save energy, enhance resource utilizations, and guarantee Quality of Service (QoS) in cloud datacenters. Most of existing studies on the virtual machine migration, however are based on a single virtual machine migration. Although there are some researches on multiple virtual machines migration, the author usually does not consider the correlation among these virtual machines. In practice, in order to save energy and maintain system performance, cloud providers usually need to migrate multiple correlated virtual machines or migrate the entire virtual datacenter (VDC) request. In this paper, we focus on the efficient online live migration of multiple correlated VMs in VDC requests, for optimizing the migration performance. To solve this problem, we propose an efficient VDC migration algorithm (VDC-M). We use the US-wide US National Science Foundation (NSF) network as substrate network to conduct extensive simulation experiments. Simulation results show that the performance of the proposed algorithm is promising in terms of the total VDC remapping cost, the blocking ratio, the average migration time and the average downtime.
Gang Sun 0001, Dan Liao, Dongcheng Zhao, Zichuan Xu, Hong-Fang Yu
IEEE Trans. Serv. Comput.3
2017 Live Migration for Service Function Chaining
Dongcheng Zhao, Gang Sun 0001, Dan Liao, Rahat Iqbal, Victor Chang 0001
IoTBDS1
2016 HMSNN: Hippocampus inspired Memory Spiking Neural Network
abstract
Human beings receive stimulations in primary sensory cortex and transfer them to higher brain regions automatically. What happened in this procedure? In this paper, we will focus on one of these regions (hippocampus) and try to simulate its working procedure by building an HMSNN (Hippocampus inspired Memory Spiking Neural Network) model. Dentate Gyrus (DG) and Cornu Ammonis area 3 (CA3) are the main regions of hippocampus and will be simulated by feed forward Spiking Neural Network (SNN) and recurrent Hopfield-like network respectively. From the structural perspective, the computational unit and the connectivity between neurons in HMSNN are all consistent with the anatomical-experimental results in hippocampus. From the functional perspective, the multi-scale memory formation, memory abstraction and memory retention will be shown in HMSNN model. In addition, the HMSNN is tested on MNIST handwritten digit dataset (with static images) and robot walking dataset (with dynamical images). The experimental result shows that: biological neural circuit inspired HMSNN shows comparable classification performance on both datasets compared to the state-of-art convolutional neural networks (CNNs), and shows significantly better performance compared to CNN when noises are introduced to the original images.
Tielin Zhang, Yi Zeng 0001, Dongcheng Zhao, Liwei Wang 0001, Yuxuan Zhao 0002, Bo Xu 0002
SMC3
2016 A new technique for efficient live migration of multiple virtual machines
Gang Sun 0001, Dan Liao, Vishal Anand 0001, Dongcheng Zhao, Hong-Fang Yu
Future Gener. Comput. Syst.4