EDBT 2026 Demo / reviewers in the wild / expert
Jibin Wu
dblp:228/1824
· DBLP profile ↗
51ranked-venue papers
7as first author
41since 2021 · last 2026
0000-0003-0135-4188ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 46 · 7 first-author · 36 since 2021Graphics, computer vision, multimedia, augmented reality and games · 13 · 1 first-author · 10 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | HardF-SNN: Hardware-Friendly Quantization for Spiking Neural Networks with Efficient Integer-Arithmetic-Only InferenceabstractSpiking Neural Networks (SNNs) are emerging as a promising energy-efficient alternative to Artificial Neural Networks (ANNs) due to their event-driven computation paradigm. However, recent advances toward large-scale high-performance SNNs inevitably lead to substantial memory and computational overhead. While quantization offers a potential way, many quantization approaches fail to deliver verifiable efficiency gains on resource-constrained hardware platforms. In this paper, we propose a lightweight and hardware-friendly SNN, termed HardF-SNN. Specifically, we first build a baseline model using shared-scale quantization and BN folding to simulate integer-only inference, as this has not been thoroughly discussed in prior SNN works. Then, through empirical and theoretical analysis, we identify that the baseline suffers from accuracy degradation and may cause training failure. To mitigate these issues, we propose proportional shared-scale quantization for enhanced dynamic range and integer-only BN using bit-shifting to stabilize training. Extensive experiments show that HardF-SNN achieves an optimal balance between performance and efficiency with excellent hardware compatibility. To demonstrate its effectiveness on resource-limited platforms, HardF-SNN is deployed on a dedicated FPGA-based hardware accelerator. Evaluation results indicate that our implementation achieves significant performance improvements over several existing hardware accelerators. Jieyuan Zhang, Yimeng Shan, Jibin Wu, Wenyu Chen 0001, Malu Zhang |
AAAI | 5 |
| 2026 | Adaptive dendritic plasticity in brain-inspired dynamic neural networks for enhanced multi-timescale feature extraction
Jiayi Mao, Hanle Zheng, Huifeng Yin, Hanxiao Fan, Lingrui Mei, Jibin Wu, Jing Pei, Lei Deng 0003 |
Neural Networks | 8 |
| 2026 | Advancing the forward-forward algorithm towards high-performance deep local learning
Yujie Wu 0002, Jibin Wu, Lei Deng 0003, Mingkun Xu, Qinghao Wen, Guoqi Li 0002 |
Neural Networks | 3 |
| 2026 | Autonomous Multiobjective Optimization Using Large Language ModelabstractMulti-objective optimization problems (MOPs) are ubiquitous in real-world applications, presenting a complex challenge of balancing multiple conflicting objectives. Traditional multi-objective evolutionary algorithms (MOEAs), though effective, often rely on domain-specific expertise for improved optimization performance, hindering adaptability to unseen MOPs. In recent years, the Large Language Models (LLMs) has revolutionized software engineering by enabling the autonomous generation and refinement of programs. Leveraging this breakthrough, we propose a new LLM-based framework that autonomously designs MOEAs for solving MOPs. The proposed framework includes a robust testing module to refine the generated MOEA through error-driven dialogue with LLMs, a dynamic selection strategy along with informative prompting-based crossover and mutation to fit textual optimization pipeline. Our approach facilitates the design of MOEA without the extensive demands for expert intervention, thereby speeding up the innovation of MOEA. Empirical studies across various MOP categories validate the robustness and superior performance of our proposed framework. Shenghao Wu, Wenjie Zhang 0004, Jibin Wu, Liang Feng 0001, Kay Chen Tan |
IEEE Trans. Evol. Comput. | 4 |
| 2026 | ISTASTrack: Bridging ANN and SNN via ISTA Adapter for RGB-Event TrackingabstractRGB-Event tracking has become a promising trend in visual object tracking to leverage the complementary strengths of both RGB images and dynamic spike events for improved performance. However, existing artificial neural networks (ANNs) struggle to fully exploit the sparse and asynchronous nature of event streams. Recent efforts toward hybrid architectures combining ANNs and spiking neural networks (SNNs) have emerged as a promising solution in RGB-Event perception, yet effectively fusing features across heterogeneous paradigms remains a challenge. In this work, we propose ISTASTrack, the first transformer-based ANN-SNN hybrid Tracker equipped with ISTA adapters for RGB-Event tracking. The two-branch model employs a vision transformer to extract spatial context from RGB inputs and a spiking transformer to capture spatio-temporal dynamics from event streams. To bridge the modality and paradigm gap between ANN and SNN features, we systematically design an ISTA adapter for bidirectional feature interaction between the two branches. The ISTA adapter is derived from the sparse representation theory by unfolding the iterative shrinkage-thresholding algorithm. Additionally, we incorporate a temporal downsampling attention module within the adapter to align multi-step SNN features with single-step ANN features in the latent space. Experimental results on RGB-Event tracking benchmarks, such as FE240hz, VisEvent, COESOT, and FELT, have demonstrated that ISTASTrack achieves state-of-the-art performance while maintaining high energy efficiency. This work highlights the effectiveness and practicality of hybrid ANN-SNN designs for robust visual tracking. The code is publicly available at https://github.com/lsying009/ISTASTrack.git. Zikai Wang 0005, Hanle Zheng, Yifan Hu 0013, Xilin Wang, Qingkai Yang, Jibin Wu, Lei Deng 0003 |
IEEE Trans. Image Process. | 7 |
| 2026 | IML-Spikeformer: Input-Aware Multilevel Spiking Transformer for Speech ProcessingabstractSpiking neural networks (SNNs), inspired by biological neural mechanisms, represent a promising neuromorphic computing paradigm that offers energy-efficient alternatives to traditional artificial neural networks (ANNs). Despite proven effectiveness, SNN architectures have struggled to achieve competitive performance on large-scale speech processing tasks. Two key challenges hinder progress: 1) the high computational overhead during training caused by multitimestep spike firing and 2) the absence of large-scale SNN architectures tailored to speech processing tasks. To overcome the issues, we introduce the input-aware multilevel spikeformer (IML-Spikeformer), a spiking transformer architecture specifically designed for large-scale speech processing. Central to our design is the input-aware multilevel spike (IMLS) mechanism, which simulates multitimestep spike firing within a single timestep using an adaptive, input-aware thresholding scheme. IML-Spikeformer further integrates a reparameterized spiking self-attention (RepSSA) module with a hierarchical decay mask (HDM), forming the HD-RepSSA module. This module enhances the precision of attention maps and enables modeling of multiscale temporal dependencies in speech signals. Experiments demonstrate that IML-Spikeformer achieves word error rates (WERs) of 6.0% on AiShell-1 and 3.4% on Librispeech-960, comparable to conventional ANN transformers while reducing theoretical inference energy consumption by $4.64\times $ and $4.32\times $ , respectively. IML-Spikeformer marks an advance of scalable SNN architectures for large-scale speech processing in both task performance and energy efficiency. Our source code and model checkpoints are publicly available at github.com/Pooookeman/IML-Spikeformer. Zeyang Song, Yuhong Chou, Jibin Wu, Haizhou Li 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2026 | SNN-FT: Temporal-Coded Spiking Neural Networks for Fourier TransformabstractThe Fourier transform (FT) stands as a fundamental tool in modern signal processing with widespread applications across various scientific and engineering fields. Therefore, there remains a need for continued research efforts to devise energy-efficient implementations of the FT. Due to their inherent energy efficiency, biologically plausible spiking neural networks (SNNs) emerge as a promising alternative solution. However, current SNN implementations of the FT suffer from two key shortcomings, namely, high latency and reduced accuracy. In this article, we analyze the underlying causes of these limitations and highlight deficiencies in the existing spike-based encoding mechanisms and spiking neuron models. We then propose a new SNN-based FT (SNN-FT) based on a logarithmically polarized time-to-first-spike (TTFS) encoding method (called LP-TTFS) along with a novel piecewise spiking neuron (PTSN) model based on ternary spikes (referred to as PTSN). The resulting SNN-FT is mathematically equivalent to the conventional FT and demonstrates superior performance in accuracy as well as reduced latency. We assess the performance of the proposed SNN-FT alternative through extensive experiments on FT-based applications, such as radar and audio signal processing, and the obtained results demonstrate the efficacy of SNN-FT and its superiority over the existing approaches. This study unveils a novel energy-efficient neuromorphic computing technique with great potential for FT applications across diverse scientific and engineering domains. Shuai Wang 0058, Haorui Zheng, Ammar Belatreche, Guoqing Wang 0001, Yeying Jin, Jibin Wu, Malu Zhang, Yang Yang 0002, Haizhou Li 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 7 |
| 2025 | Towards Robustness and Explainability of Automatic Algorithm SelectionabstractAlgorithm selection aims to identify the optimal performing algorithm before execution. Existing techniques typically focus on the observed correlations between algorithm performance and meta-features. However, little research has explored the underlying mechanisms of algorithm selection, specifically what characteristics an algorithm must possess to effectively tackle problems with certain feature values. This gap not only limits the explainability but also makes existing models vulnerable to data bias and distribution shift. This paper introduces directed acyclic graph (DAG) to describe this mechanism, proposing a novel modeling paradigm that aligns more closely with the fundamental logic of algorithm selection. By leveraging DAG to characterize the algorithm feature distribution conditioned on problem features, our approach enhances robustness against marginal distribution changes and allows for finer-grained predictions through the reconstruction of optimal algorithm features, with the final decision relying on differences between reconstructed and rejected algorithm features. Furthermore, we demonstrate that, the learned DAG and the proposed counterfactual calculations offer our approach with both model-level and instance-level explainability. Jibin Wu, Yu Zhou 0045, Liang Feng 0001, KC Tan |
ICML | 2 |
| 2025 | KoopSTD: Reliable Similarity Analysis between Dynamical Systems via Approximating Koopman Spectrum with Timescale DecouplingabstractDetermining the similarity between dynamical systems remains a long-standing challenge in both machine learning and neuroscience. Recent works based on Koopman operator theory have proven effective in analyzing dynamical similarity by examining discrepancies in the Koopman spectrum. Nevertheless, existing similarity metrics can be severely constrained when systems exhibit complex nonlinear behaviors across multiple temporal scales. In this work, we propose KoopSTD, a dynamical similarity measurement framework that precisely characterizes the underlying dynamics by approximating the Koopman spectrum with explicit timescale decoupling and spectral residual control. We show that KoopSTD maintains invariance under several common representation-space transformations, which ensures robust measurements across different coordinate systems. Our extensive experiments on physical and neural systems validate the effectiveness, scalability, and robustness of KoopSTD compared to existing similarity metrics. We also apply KoopSTD to explore two open-ended research questions in neuroscience and large language models, highlighting its potential to facilitate future scientific and engineering discoveries. Code is available at link. Ziyuan Ye, Yinsong Yan, Zeyang Song, Yujie Wu 0002, Jibin Wu |
ICML | 6 |
| 2025 | Neuromorphic Sequential Arena: A Benchmark for Neuromorphic Temporal ProcessingabstractTemporal processing is vital for extracting meaningful information from time-varying signals. Recent advancements in Spiking Neural Networks (SNNs) have shown immense promise in efficiently processing these signals. However, progress in this field has been impeded by the lack of effective and standardized benchmarks, which complicates the consistent measurement of technological advancements and limits the practical applicability of SNNs. To bridge this gap, we introduce the Neuromorphic Sequential Arena (NSA), a comprehensive benchmark that offers an effective, versatile, and application-oriented evaluation framework for neuromorphic temporal processing. The NSA includes seven real-world temporal processing tasks from a diverse range of application scenarios, each capturing rich temporal dynamics across multiple timescales. Utilizing NSA, we conduct extensive comparisons of recently introduced spiking neuron models and neural architectures, presenting comprehensive baselines in terms of task performance, training speed, memory usage, and energy efficiency. Our findings emphasize an urgent need for efficient SNN designs that can consistently deliver high performance across tasks with varying temporal complexities while maintaining low computational costs. NSA enables systematic tracking of advancements in neuromorphic algorithm research and paves the way for developing effective and efficient neuromorphic temporal processing systems. Chenxiang Ma, Yujie Wu 0002, Kay Chen Tan, Jibin Wu |
IJCAI | 5 |
| 2025 | MSVIT: Improving Spiking Vision Transformer Using Multi-scale Attention FusionabstractThe combination of Spiking Neural Networks (SNNs) with Vision Transformer architectures has attracted significant attention due to the great potential for energy-efficient and high-performance computing paradigms. However, a substantial performance gap still exists between SNN-based and ANN-based transformer architectures. While existing methods propose spiking self-attention mechanisms that are successfully combined with SNNs, the overall architectures proposed by these methods suffer from a bottleneck in effectively extracting features from different image scales. In this paper, we address this issue and propose MSVIT, a novel spike-driven Transformer architecture, which firstly uses multi-scale spiking attention (MSSA) to enrich the capability of spiking attention blocks. We validate our approach across various main data sets. The experimental results indicate that our MSVIT outperforms existing SNN-based models, positioning itself as a state-of-the-art solution among NN-transformer architectures. The codes are available at https://github.com/Nanhu-AI-Lab/MSViT. Chenlin Zhou, Jibin Wu, Yansong Chua, Yangyang Shu |
IJCAI | 3 |
| 2025 | ZeCO: Zero-Communication Overhead Sequence Parallelism for Linear AttentionabstractLinear attention mechanisms deliver significant advantages for Large Language Models (LLMs) by providing linear computational complexity, enabling efficient processing of ultra-long sequences (e.g., 1M context). However, existing Sequence Parallelism (SP) methods, essential for distributing these workloads across devices, become the primary performance bottleneck due to substantial communication overhead. In this paper, we introduce ZeCO (Zero Communication Overhead) sequence parallelism for linear attention models, a new SP method designed to overcome these limitations and achieve practically end-to-end near-linear scalability for long sequence training. For example, training a model with a 1M sequence length across 64 devices using ZeCO takes roughly the same time as training with an 16k sequence on a single device. At the heart of ZeCO lies All-Scan, a novel collective communication primitive. All-Scan provides each SP rank with precisely the initial operator state it requires while maintaining a minimal communication footprint, effectively eliminating communication overhead. Theoretically, we prove the optimaity of ZeCO, showing that it introduces only negligible time and space overhead. Empirically, we compare the communication costs of different sequence parallelism strategies and demonstrate that All-Scan achieves the fastest communication in SP scenarios. Specifically, on 256 GPUs with an 8M sequence length, ZeCO achieves a 60\% speedup compared to the current state-of-the-art (SOTA) SP method. We believe ZeCO establishes a clear path toward efficiently training next-generation LLMs on previously intractable sequence lengths. Yuhong Chou, Rui-Jie Zhu 0003, Tianjian Li, Congying Chu, Qian Liu 0033, Jibin Wu, Zejun Ma 0001 |
NeurIPS | 8 |
| 2025 | Diversity-Aware Policy Optimization for Large Language Model ReasoningabstractThe reasoning capabilities of large language models (LLMs) have advanced rapidly, particularly following the release of DeepSeek-R1, which has inspired a surge of research into data quality and reinforcement learning (RL) algorithms. Despite the pivotal role diversity plays in RL, its influence on LLM reasoning remains largely underexplored. To bridge this gap, this work presents a systematic investigation into the impact of diversity in RL-based training for LLM reasoning, and proposes a novel diversity-aware policy optimization method. Across evaluations on 12 LLMs, we observe a strong positive correlation between the solution diversity and potential@k (a novel metric quantifying an LLM’s reasoning potential) in high-performing models. This finding motivates our method to explicitly promote diversity during RL training. Specifically, we design a token-level diversity and reformulate it into a practical objective, then we selectively apply it to positive samples. Integrated into the R1-zero training framework, our method achieves a 3.5\% average improvement across four mathematical reasoning benchmarks, while generating more diverse and robust solutions. Ran Cheng 0004, Jibin Wu, KC Tan |
NeurIPS | 4 |
| 2025 | HM3: Hierarchical Multi-Objective Model Merging for Pretrained ModelsabstractModel merging is a technique that combines multiple large pretrained models into a single model, enhancing performance and broadening task adaptability without original data or additional training. However, most existing model merging methods focus primarily on exploring the parameter space, merging models with identical architectures. Despite its potential, merging in the architecture space remains in its early stages due to the vast search space and challenges related to layer compatibility. This paper designs a hierarchical model merging framework named HM3, formulating a bilevel multi-objective model merging problem across both parameter and architecture spaces. At the parameter level, HM3 integrates existing merging methods to quickly identify optimal parameters. Based on these, an actor-critic strategy with efficient policy discretization is employed at the architecture level to explore inference paths with Markov property in the layer-granularity search space for reconstructing these optimal models. By training reusable policy and value networks, HM3 learns Pareto optimal models to provide customized solutions for various tasks. Experimental results on language and vision tasks demonstrate that HM3 outperforms methods focusing solely on the parameter or architecture space. Yu Zhou 0045, Jibin Wu, Liang Feng 0001, KC Tan |
NeurIPS | 3 |
| 2025 | CausalMixNet: A mixed-attention framework for causal intervention in robust medical image diagnosis
Yao Hu 0001, Rui Liu 0038, Jibin Wu, Zhi-an Huang, Kay Chen Tan |
Medical Image Anal. | 5 |
| 2025 | Towards parameter-free attentional spiking neural networks
Pengfei Sun 0003, Jibin Wu, Paul Devos, Dick Botteldooren |
Neural Networks | 2 |
| 2025 | A Fine-grained Hemispheric Asymmetry Network for accurate and interpretable EEG-based emotion classificationabstractIn this work, we propose a Fine-grained Hemispheric Asymmetry Network (FG-HANet), an end-to-end deep learning model that leverages hemispheric asymmetry features within 2-Hz narrow frequency bands for accurate and interpretable emotion classification over raw EEG data. In particular, the FG-HANet extracts features not only from original inputs but also from their mirrored versions, and applies Finite Impulse Response (FIR) filters at a granularity as fine as 2-Hz to acquire fine-grained spectral information. Furthermore, to guarantee sufficient attention to hemispheric asymmetry features, we tailor a three-stage training pipeline for the FG-HANet to further boost its performance. We conduct extensive evaluations on two public datasets, SEED and SEED-IV, and experimental results well demonstrate the superior performance of the proposed FG-HANet, i.e. 97.11% and 85.70% accuracy, respectively, building a new state-of-the-art. Our results also reveal the hemispheric dominance under different emotional states and the hemisphere asymmetry within 2-Hz frequency bands in individuals. These not only align with previous findings in neuroscience but also provide new insights into underlying emotion generation mechanisms. Ruofan Yan, Yuxuan Yan, Xu Niu, Jibin Wu |
Neural Networks | 5 |
| 2025 | Dynamic Graph Representation Learning for Spatio-Temporal Neuroimaging AnalysisabstractNeuroimaging analysis aims to reveal the information-processing mechanisms of the human brain in a noninvasive manner. In the past, graph neural networks (GNNs) have shown promise in capturing the non-Euclidean structure of brain networks. However, existing neuroimaging studies focused primarily on spatial functional connectivity, despite temporal dynamics in complex brain networks. To address this gap, we propose a spatio-temporal interactive graph representation framework (STIGR) for dynamic neuroimaging analysis that encompasses different aspects from classification and regression tasks to interpretation tasks. STIGR leverages a dynamic adaptive-neighbor graph convolution network to capture the interrelationships between spatial and temporal dynamics. To address the limited global scope in graph convolutions, a self-attention module based on Transformers is introduced to extract long-term dependencies. Contrastive learning is used to adaptively contrast similarities between adjacent scanning windows, modeling cross-temporal correlations in dynamic graphs. Extensive experiments on six public neuroimaging datasets demonstrate the competitive performance of STIGR across different platforms, achieving state-of-the-art results in classification and regression tasks. The proposed framework enables the detection of remarkable temporal association patterns between regions of interest based on sequential neuroimaging signals, offering medical professionals a versatile and interpretable tool for exploring task-specific neurological patterns. Our codes and models are available at https://github.com/77YQ77/STIGR/. Rui Liu 0038, Yao Hu 0001, Jibin Wu, Ka-Chun Wong, Zhi-an Huang, Kay Chen Tan |
IEEE Trans. Cybern. | 3 |
| 2025 | Evolutionary Computation in the Era of Large Language Model: Survey and RoadmapabstractLarge language models (LLMs) have not only revolutionized natural language processing but also extended their prowess to various domains, marking a significant stride toward artificial general intelligence. The interplay between LLMs and evolutionary algorithms (EAs), despite differing in objectives and methodologies, share a common pursuit of applicability in complex problems. Meanwhile, EA can provide an optimization framework for LLM’s further enhancement under closed box settings, empowering LLM with flexible global search capacities. On the other hand, the abundant domain knowledge inherent in LLMs could enable EA to conduct more intelligent searches. Furthermore, the text processing and generative capabilities of LLMs would aid in deploying EAs across a wide range of tasks. Based on these complementary advantages, this article provides a thorough review and a forward-looking roadmap, categorizing the reciprocal inspiration into two main avenues: 1) LLM-enhanced EA and 2) EA-enhanced LLM. Some integrated synergy methods are further introduced to exemplify the complementarity between LLMs and EAs in diverse scenarios, including code generation, software engineering, neural architecture search, and various generation tasks. As the first comprehensive review focused on the EA research in the era of LLMs, this article provides a foundational stepping stone for understanding the collaborative potential of LLMs and EAs. The identified challenges and future directions offer guidance for researchers and practitioners to unlock the full potential of this innovative collaboration in propelling advancements in optimization and artificial intelligence. We have created a GitHub repository to index the relevant papers:https://github.com/wuxingyu-ai/LLM4EC. Sheng-Hao Wu, Jibin Wu, Liang Feng 0001, Kay Chen Tan |
IEEE Trans. Evol. Comput. | 3 |
| 2025 | Toward Ultralow-Power Neuromorphic Speech Enhancement With Spiking-FullSubNetabstractSpeech enhancement (SE) is critical for improving speech intelligibility and quality in various audio devices. In recent years, deep learning-based methods have significantly improved SE performance, but they often come with a high computational cost, which is prohibitive for a large number of edge devices, such as headsets and hearing aids. This work proposes an ultralow-power SE system based on the brain-inspired spiking neural network (SNN) called Spiking-FullSubNet. Spiking-FullSubNet follows a full-band and subband fusioned approach to effectively capture both global and local spectral information. To enhance the efficiency of computationally expensive subband modeling, we introduce a frequency partitioning method inspired by the sensitivity profile of the human peripheral auditory system. Furthermore, we introduce a novel spiking neuron model that can dynamically control the input information integration and forgetting, enhancing the multiscale temporal processing capability of SNN, which is critical for speech denoising. Experiments conducted on the recent Intel Neuromorphic Deep Noise Suppression (N-DNS) Challenge dataset show that the Spiking-FullSubNet surpasses state-of-the-art (SOTA) methods by large margins in terms of both speech quality and energy efficiency metrics. Notably, our system won the championship of the Intel N-DNS Challenge (algorithmic track), opening up a myriad of opportunities for ultralow-power SE at the edge. Our source code and model checkpoints are publicly available at github.com/haoxiangsnr/spiking-fullsubnet. Chenxiang Ma, Qu Yang, Jibin Wu, Kay Chen Tan |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2025 | Delayed Memory Unit: Modeling Temporal Dependency Through Delay GateabstractRecurrent neural networks (RNNs) are widely recognized for their proficiency in modeling temporal dependencies, making them highly prevalent in sequential data processing applications. Nevertheless, vanilla RNNs are confronted with the well-known issue of gradient vanishing and exploding, posing a significant challenge for learning and establishing long-range dependencies. Additionally, gated RNNs tend to be over-parameterized, resulting in poor computational efficiency and network generalization. To address these challenges, this article proposes a novel delayed memory unit (DMU). The DMU incorporates a delay line structure along with delay gates into vanilla RNN, thereby enhancing temporal interaction and facilitating temporal credit assignment. Specifically, the DMU is designed to directly distribute the input information to the optimal time instant in the future, rather than aggregating and redistributing it over time through intricate network dynamics. Our proposed DMU demonstrates superior temporal modeling capabilities across a broad range of sequential modeling tasks, utilizing considerably fewer parameters than other state-of-the-art gated RNN models in applications such as speech recognition, radar gesture recognition, ECG waveform segmentation, and permuted sequential (PS) image classification. Pengfei Sun 0003, Jibin Wu, Malu Zhang, Paul Devos, Dick Botteldooren |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2025 | Toward Building Human-Like Sequential Memory Using Brain-Inspired Spiking Neural ModelsabstractThe brain is able to acquire and store memories of everyday experiences in real-time. It can also selectively forget information to facilitate memory updating. However, our understanding of the underlying mechanisms and coordination of these processes within the brain remains limited. However, no existing artificial intelligence models have yet matched human-level capabilities in terms of memory storage and retrieval. This study introduces a brain-inspired spiking neural model that integrates the learning and forgetting processes of sequential memory. The proposed model closely mimics the distributed and sparse temporal coding observed in the biological neural system. It employs one-shot online learning for memory formation and uses biologically plausible mechanisms of neural oscillation and phase precession to retrieve memorized sequences reliably. In addition, an active forgetting mechanism is integrated into the spiking neural model, enabling memory removal, flexibility, and updating. The proposed memory model not only enhances our understanding of human memory processes but also provides a robust framework for addressing temporal modeling tasks. Malu Zhang, Xiaoling Luo 0001, Jibin Wu, Ammar Belatreche, Siqi Cai 0002, Yang Yang 0002, Haizhou Li 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2024 | TC-LIF: A Two-Compartment Spiking Neuron Model for Long-Term Sequential ModellingabstractThe identification of sensory cues associated with potential opportunities and dangers is frequently complicated by unrelated events that separate useful cues by long delays. As a result, it remains a challenging task for state-of-the-art spiking neural networks (SNNs) to establish long-term temporal dependency between distant cues. To address this challenge, we propose a novel biologically inspired Two-Compartment Leaky Integrate-and-Fire spiking neuron model, dubbed TC-LIF. The proposed model incorporates carefully designed somatic and dendritic compartments that are tailored to facilitate learning long-term temporal dependencies. Furthermore, the theoretical analysis is provided to validate the effectiveness of TC-LIF in propagating error gradients over an extended temporal duration. Our experimental results, on a diverse range of temporal classification tasks, demonstrate superior temporal classification capability, rapid training convergence, and high energy efficiency of the proposed TC-LIF model. Therefore, this work opens up a myriad of opportunities for solving challenging temporal processing tasks on emerging neuromorphic computing systems. Our code is publicly available at https://github.com/ZhangShimin1/TC-LIF. Qu Yang, Chenxiang Ma, Jibin Wu, Haizhou Li 0001, Kay Chen Tan |
AAAI | 4 |
| 2024 | Spiking-Leaf: A Learnable Auditory Front-End for Spiking Neural NetworksabstractBrain-inspired spiking neural networks (SNNs) have demonstrated great potential for temporal signal processing. However, their performance in speech processing remains limited due to the lack of an effective auditory front-end. To address this limitation, we introduce Spiking-LEAF, a learnable auditory front-end meticulously designed for SNN-based speech processing. Spiking-LEAF combines a learnable filter bank with a novel two-compartment spiking neuron model called IHC-LIF. The IHC-LIF neurons draw inspiration from the structure of inner hair cells (IHC) and they leverage segregated dendritic and somatic compartments to effectively capture multi-scale temporal dynamics of speech signals. Additionally, the IHC-LIF neurons incorporate the lateral feedback mechanism along with spike regularization loss to enhance spike encoding efficiency. On keyword spotting and speaker identification tasks, the proposed Spiking-LEAF outperforms both SOTA spiking auditory front-ends and conventional real-valued acoustic features in terms of classification accuracy, noise robustness, and encoding efficiency. Zeyang Song, Jibin Wu, Malu Zhang, Zheng Shou 0001, Haizhou Li 0001 |
ICASSP | 2 |
| 2024 | Scaling Supervised Local Learning with Augmented Auxiliary NetworksabstractDeep neural networks are typically trained using global error signals that backpropagate (BP) end-to-end, which is not only biologically implausible but also suffers from the update locking problem and requires huge memory consumption. Local learning, which updates each layer independently with a gradient-isolated auxiliary network, offers a promising alternative to address the above problems. However, existing local learning methods are confronted with a large accuracy gap with the BP counterpart, particularly for large-scale networks. This is due to the weak coupling between local layers and their subsequent network layers, as there is no gradient communication across layers. To tackle this issue, we put forward an augmented local learning method, dubbed AugLocal. AugLocal constructs each hidden layer’s auxiliary network by uniformly selecting a small subset of layers from its subsequent network layers to enhance their synergy. We also propose to linearly reduce the depth of auxiliary networks as the hidden layer goes deeper, ensuring sufficient network capacity while reducing the computational cost of auxiliary networks. Our extensive experiments on four image classification datasets (i.e., CIFAR-10, SVHN, STL-10, and ImageNet) demonstrate that AugLocal can effectively scale up to tens of local layers with a comparable accuracy to BP-trained networks while reducing GPU memory usage by around 40%. The proposed AugLocal method, therefore, opens up a myriad of opportunities for training high-performance deep neural networks on resource-constrained platforms. Code is available at \url{https://github.com/ChenxiangMA/AugLocal}. Chenxiang Ma, Jibin Wu, Chenyang Si, Kay Chen Tan |
ICLR | 2 |
| 2024 | Large Language Model-Enhanced Algorithm Selection: Towards Comprehensive Algorithm Representation
Yan Zhong 0001, Jibin Wu, Bingbing Jiang 0001, Kay Chen Tan |
IJCAI | 3 |
| 2024 | Efficient Online Learning for Networks of Two-Compartment Spiking NeuronsabstractThe brain-inspired Spiking Neural Networks (SNNs) have garnered considerable research interest due to their superior performance and energy efficiency in processing temporal signals. Recently, a novel multi-compartment spiking neuron model, namely the Two-Compartment LIF (TC-LIF) model, has been proposed and exhibited a remarkable capacity for sequential modelling. However, training the TC-LIF model presents challenges stemming from the large memory consumption and the issue of vanishing gradient associated with the Backpropagation Through Time (BPTT) algorithm. To address these challenges, online learning methodologies emerge as a promising solution. Yet, to date, the application of online learning methods in SNNs has been predominantly confined to simplified Leaky Integrate-and-Fire (LIF) neuron models. In this paper, we present a novel online learning method specifically tailored for networks of TC-LIF neurons. Additionally, we propose a refined TC-LIF neuron model called Adaptive TC-LIF, which is carefully designed to enhance temporal information integration in online learning scenarios. Extensive experiments, conducted on various sequential benchmarks, demonstrate that our approach successfully preserves the superior sequential modeling capabilities of the TC-LIF neuron while incorporating the training efficiency and hardware friendliness of online learning. As a result, it offers a multitude of opportunities to leverage neuromorphic solutions for processing temporal signals. Yujia Yin, Chenxiang Ma, Jibin Wu, Kay Chen Tan |
IJCNN | 4 |
| 2024 | Global-Local Convolution with Spiking Neural Networks for Energy-efficient Keyword Spottingabstract25th Annual Conference of the International Speech Communication Association, Interspeech 2024, Kos, Greece, September 1-5, 2024. ISCA 2024 Shuai Wang 0058, Dehao Zhang, Wenjie Wei, Jibin Wu, Malu Zhang |
INTERSPEECH | 6 |
| 2024 | Mixed Prototype Correction for Causal Inference in Medical Image ClassificationabstractThe heterogeneity of medical images poses significant challenges to accurate disease diagnosis. To tackle this issue, the impact of such heterogeneity on the causal relationship between image features and diagnostic labels should be incorporated into model design, which however remains underexplored. In this paper, we propose a mixed prototype correction for causal inference (MPCCI) method, aimed at mitigating the impact of unseen confounding factors on the causal relationships between medical images and disease labels, so as to enhance the diagnostic accuracy of deep learning models. The MPCCI comprises a causal inference component based on front-door adjustment and an adaptive training strategy. The causal inference component employs a multi-view feature extraction (MVFE) module to establish mediators, and a mixed prototype correction (MPC) module to execute causal interventions. Moreover, the adaptive training strategy incorporates both information purity and maturity metrics to maintain stable model training. Experimental evaluations on four medical image datasets, encompassing CT and ultrasound modalities, demonstrate the superior diagnostic accuracy and reliability of the proposed MPCCI. The code will be available at https://github.com/Yajie-Zhang/MPCCI. Zhi-an Huang, Zhiliang Hong 0002, Songsong Wu, Jibin Wu, Kay Chen Tan |
ACM Multimedia | 5 |
| 2024 | MetaLA: Unified Optimal Linear Approximation to Softmax Attention MapabstractVarious linear complexity models, such as Linear Transformer (LinFormer), State Space Model (SSM), and Linear RNN (LinRNN), have been proposed to replace the conventional softmax attention in Transformer structures. However, the optimal design of these linear models is still an open question. In this work, we attempt to answer this question by finding the best linear approximation to softmax attention from a theoretical perspective. We start by unifying existing linear complexity models as the linear attention form and then identify three conditions for the optimal linear attention design: (1) Dynamic memory ability; (2) Static approximation ability; (3) Least parameter approximation. We find that none of the current linear models meet all three conditions, resulting in suboptimal performance. Instead, we propose Meta Linear Attention (MetaLA) as a solution that satisfies these conditions. Our experiments on Multi-Query Associative Recall (MQAR) task, language modeling, image classification, and Long-Range Arena (LRA) benchmark demonstrate that MetaLA is more effective than the existing linear models. Yuhong Chou, Man Yao, Yuqi Pan, Rui-Jie Zhu 0003, Jibin Wu, Yiran Zhong, Bo Xu 0002, Guoqi Li 0002 |
NeurIPS | 6 |
| 2024 | Delay learning based on temporal coding in Spiking Neural Networks
Pengfei Sun 0003, Jibin Wu, Malu Zhang, Paul Devos, Dick Botteldooren |
Neural Networks | 2 |
| 2024 | A Hybrid Neural Coding Approach for Pattern Recognition With Spiking Neural NetworksabstractRecently, brain-inspired spiking neural networks (SNNs) have demonstrated promising capabilities in solving pattern recognition tasks. However, these SNNs are grounded on homogeneous neurons that utilize a uniform neural coding for information representation. Given that each neural coding scheme possesses its own merits and drawbacks, these SNNs encounter challenges in achieving optimal performance such as accuracy, response time, efficiency, and robustness, all of which are crucial for practical applications. In this study, we argue that SNN architectures should be holistically designed to incorporate heterogeneous coding schemes. As an initial exploration in this direction, we propose a hybrid neural coding and learning framework, which encompasses a neural coding zoo with diverse neural coding schemes discovered in neuroscience. Additionally, it incorporates a flexible neural coding assignment strategy to accommodate task-specific requirements, along with novel layer-wise learning methods to effectively implement hybrid coding SNNs. We demonstrate the superiority of the proposed framework on image classification and sound localization tasks. Specifically, the proposed hybrid coding SNNs achieve comparable accuracy to state-of-the-art SNNs, while exhibiting significantly reduced inference latency and energy consumption, as well as high noise robustness. This study yields valuable insights into hybrid neural coding designs, paving the way for developing high-performance neuromorphic systems. Qu Yang, Jibin Wu, Haizhou Li 0001, Kay Chen Tan |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2023 | A Tandem Learning Rule for Effective Training and Rapid Inference of Deep Spiking Neural NetworksabstractSpiking neural networks (SNNs) represent the most prominent biologically inspired computing model for neuromorphic computing (NC) architectures. However, due to the nondifferentiable nature of spiking neuronal functions, the standard error backpropagation algorithm is not directly applicable to SNNs. In this work, we propose a tandem learning framework that consists of an SNN and an artificial neural network (ANN) coupled through weight sharing. The ANN is an auxiliary structure that facilitates the error backpropagation for the training of the SNN at the spike-train level. To this end, we consider the spike count as the discrete neural representation in the SNN and design an ANN neuronal activation function that can effectively approximate the spike count of the coupled SNN. The proposed tandem learning rule demonstrates competitive pattern recognition and regression capabilities on both the conventional frame- and event-based vision datasets, with at least an order of magnitude reduced inference time and total synaptic operations over other state-of-the-art SNN implementations. Therefore, the proposed tandem learning rule offers a novel solution to training efficient, low latency, and high-accuracy deep SNNs with low computing resources. Jibin Wu, Yansong Chua, Malu Zhang, Guoqi Li 0002, Haizhou Li 0001, Kay Chen Tan |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2022 | A Hybrid Learning Framework for Deep Spiking Neural Networks with One-Spike Temporal CodingabstractBio-inspired spiking neural networks (SNNs) are compelling candidates for spatio-temporal information processing on ultra-low power neuromorphic computing chips. However, the existing SNN training methods have not fully exploited the temporal information of spikes that plays a critical role in sparse information representation and communication. Hereby, we present a hybrid learning framework for deep SNNs with one-spike temporal coding to make full utilization of the spike timing. We first propose a novel ANN-to-SNN conversion method based on forward propagation mechanisms of ANNs and SNNs to offer a good initialization for SNNs. The performance of the converted SNN is further improved by training with a timing-based backpropagation (BP) method. Experimental results demonstrate that the proposed hybrid learning framework can achieve competitive accuracies on both visual and audio recognition tasks with significantly improved training efficiency over direct SNN BP methods. Jibin Wu, Malu Zhang, Qi Liu 0005, Haizhou Li 0001 |
ICASSP | 2 |
| 2022 | Training Spiking Neural Networks with Local Tandem LearningabstractSpiking neural networks (SNNs) are shown to be more biologically plausible and energy efficient over their predecessors. However, there is a lack of an efficient and generalized training method for deep SNNs, especially for deployment on analog computing substrates. In this paper, we put forward a generalized learning rule, termed Local Tandem Learning (LTL). The LTL rule follows the teacher-student learning approach by mimicking the intermediate feature representations of a pre-trained ANN. By decoupling the learning of network layers and leveraging highly informative supervisor signals, we demonstrate rapid network convergence within five training epochs on the CIFAR-10 dataset while having low computational complexity. Our experimental results have also shown that the SNNs thus trained can achieve comparable accuracies to their teacher ANNs on CIFAR-10, CIFAR-100, and Tiny ImageNet datasets. Moreover, the proposed LTL rule is hardware friendly. It can be easily implemented on-chip to perform fast parameter calibration and provide robustness against the notorious device non-ideality issues. It, therefore, opens up a myriad of opportunities for training and deployment of SNN on ultra-low-power mixed-signal neuromorphic computing chips. Qu Yang, Jibin Wu, Malu Zhang, Yansong Chua, Xinchao Wang, Haizhou Li 0001 |
NeurIPS | 2 |
| 2022 | Progressive Tandem Learning for Pattern Recognition With Deep Spiking Neural NetworksabstractSpiking neural networks (SNNs) have shown clear advantages over traditional artificial neural networks (ANNs) for low latency and high computational efficiency, due to their event-driven nature and sparse communication. However, the training of deep SNNs is not straightforward. In this paper, we propose a novel ANN-to-SNN conversion and layer-wise learning framework for rapid and efficient pattern recognition, which is referred to as progressive tandem learning. By studying the equivalence between ANNs and SNNs in the discrete representation space, a primitive network conversion method is introduced that takes full advantage of spike count to approximate the activation value of ANN neurons. To compensate for the approximation errors arising from the primitive network conversion, we further introduce a layer-wise learning method with an adaptive training scheduler to fine-tune the network weights. The progressive tandem learning framework also allows hardware constraints, such as limited weight precision and fan-in connections, to be progressively imposed during training. The SNNs thus trained have demonstrated remarkable classification and regression capabilities on large-scale object recognition, image reconstruction, and speech separation tasks, while requiring at least an order of magnitude reduced inference time and synaptic operations than other state-of-the-art SNN implementations. It, therefore, opens up a myriad of opportunities for pervasive mobile and embedded devices with a limited power budget. Jibin Wu, Chenglin Xu, Daquan Zhou, Malu Zhang, Haizhou Li 0001, Kay Chen Tan |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2022 | Rectified Linear Postsynaptic Potential Function for Backpropagation in Deep Spiking Neural NetworksabstractSpiking neural networks (SNNs) use spatiotemporal spike patterns to represent and transmit information, which are not only biologically realistic but also suitable for ultralow-power event-driven neuromorphic implementation. Just like other deep learning techniques, deep SNNs (DeepSNNs) benefit from the deep architecture. However, the training of DeepSNNs is not straightforward because the well-studied error backpropagation (BP) algorithm is not directly applicable. In this article, we first establish an understanding as to why error BP does not work well in DeepSNNs. We then propose a simple yet efficient rectified linear postsynaptic potential function (ReL-PSP) for spiking neurons and a spike-timing-dependent BP (STDBP) learning algorithm for DeepSNNs where the timing of individual spikes is used to convey information (temporal coding), and learning (BP) is performed based on spike timing in an event-driven manner. We show that DeepSNNs trained with the proposed single spike time-based learning algorithm can achieve the state-of-the-art classification accuracy. Furthermore, by utilizing the trained model parameters obtained from the proposed STDBP learning algorithm, we demonstrate ultralow-power inference operations on a recently proposed neuromorphic inference accelerator. The experimental results also show that the neuromorphic hardware consumes 0.751 mW of the total power consumption and achieves a low latency of 47.71 ms to classify an image from the Modified National Institute of Standards and Technology (MNIST) dataset. Overall, this work investigates the contribution of spike timing dynamics for information encoding, synaptic plasticity, and decision-making, providing a new perspective to the design of future DeepSNNs and neuromorphic hardware. Malu Zhang, Jibin Wu, Ammar Belatreche, Burin Amornpaisannon, Venkata Pavan Kumar Miriyala, Hong Qu 0002, Yansong Chua, Trevor E. Carlson, Haizhou Li 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2021 | Rethinking Benchmarks for Neuromorphic Learning AlgorithmsabstractWe rely on benchmarking datasets to monitor the research progress. However, recent studies have cast doubts on the effectiveness of current neuromorphic benchmarking datasets; and the debate remains largely unsettled. In this paper, we assess the richness and usefulness of temporal information embedded in these benchmarking datasets for SNN decision making. To this end, we propose a segregated spatio-temporal learning framework that allows us to selectively control the information flow along both spatial and temporal directions during feedforward and backward propagation. Leveraging on this framework, we conduct a comprehensive study on seven widely used neuromorphic audio and vision datasets. Our findings are threefold. First, the existing neuromorphic benchmarks only make limited contributions in highlighting the temporal processing capability of spiking neurons. Second, the temporal credit assignment is redundant for tasks that only require short-range temporal dependency. Third, we recommend the neuromorphic research community to develop novel benchmarks that require both short-range and long-range temporal dependencies. Such appropriate benchmark datasets would be helpful in guiding the development of powerful SNN-based learning algorithms and computational models. Qu Yang, Jibin Wu, Haizhou Li 0001 |
IJCNN | 2 |
| 2021 | HuRAI: A brain-inspired computational model for human-robot auditory interface
Jibin Wu, Qi Liu 0005, Malu Zhang, Zihan Pan, Haizhou Li 0001, Kay Chen Tan |
Neurocomputing | 1 |
| 2021 | Multi-Tone Phase Coding of Interaural Time Difference for Sound Source Localization With Spiking Neural NetworksabstractMammals exhibit remarkable capability of detecting and localizing sound sources in complex acoustic environments by using binaural cues in the spiking manner. Emulating the auditory process for sound source localization (SSL) by mammals, we propose a computational model for accurate and robust SSL under the neuromorphic spiking neural network (SNN) framework. The center of this model is a Multi-Tone Phase Coding (MTPC) scheme, which encodes the interaural time difference (ITD) between binaural pure tones into discriminative spike patterns that can be directly classified by SNNs. As such, SSL can be implemented as an event-driven task on highly efficient, neuromorphic parallel processors. We evaluate the proposed computational model on a directional audio dataset recorded from a microphone array in a realistic acoustic environment with background noise, obstruction, reflection, and other interferences. We report superior localization capability with a mean absolute error (MAE) of 1.02°or 100% classification accuracy with an angle resolution of 5°, which surpasses other SNN-based biologically plausible neuromorphic approaches by a relatively large margin and on par with human performance in similar tasks. This study opens up many application opportunities in human-robot interaction where energy efficiency is crucial. As a case study, we successfully deploy the proposed SSL system in a robotic platform to track the speaker and orient the robot's attention. Zihan Pan, Malu Zhang, Jibin Wu, Haizhou Li 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2021 | Target Speaker Verification With Selective Auditory Attention for Single and Multi-Talker SpeechabstractSpeaker verification has been studied mostly under the single-talker condition. It is adversely affected in the presence of interference speakers. Inspired by the study on target speaker extraction, e.g., SpEx, we propose a unified speaker verification framework for both single- and multi-talker speech, that is able to pay selective auditory attention to the target speaker. This target speaker verification (tSV) framework jointly optimizes a speaker attention module and a speaker representation module via multi-task learning. We study four different target speaker embedding schemes under the tSV framework. The experimental results show that all four target speaker embedding schemes significantly outperform other competitive solutions for multi-talker speech. Notably, the best tSV speaker embedding scheme achieves 76.0% and 55.3% relative improvements over the baseline system on the WSJ0-2mix-extr and Libri2Mix corpora in terms of equal-error-rate for 2-talker speech, while the performance of tSV for single-talker speech is on par with that of traditional speaker verification system, that is trained and evaluated under the same single-talker condition. Chenglin Xu, Wei Rao 0002, Jibin Wu, Haizhou Li 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2020 | Deep Convolutional Spiking Neural Networks for Keyword Spotting
Emre Yilmaz 0001, Özgür Bora Gevrek, Jibin Wu, Xuanbo Meng, Haizhou Li 0001 |
INTERSPEECH | 3 |
| 2020 | Supervised learning in spiking neural networks with synaptic delay-weight plasticity
Malu Zhang, Jibin Wu, Ammar Belatreche, Zihan Pan, Xiurui Xie, Yansong Chua, Guoqi Li 0002, Hong Qu 0002, Haizhou Li 0001 |
Neurocomputing | 2 |
| 2019 | MPD-AL: An Efficient Membrane Potential Driven Aggregate-Label Learning Algorithm for Spiking NeuronsabstractOne of the long-standing questions in biology and machine learning is how neural networks may learn important features from the input activities with a delayed feedback, commonly known as the temporal credit-assignment problem. The aggregate-label learning is proposed to resolve this problem by matching the spike count of a neuron with the magnitude of a feedback signal. However, the existing threshold-driven aggregate-label learning algorithms are computationally intensive, resulting in relatively low learning efficiency hence limiting their usability in practical applications. In order to address these limitations, we propose a novel membrane-potential driven aggregate-label learning algorithm, namely MPD-AL. With this algorithm, the easiest modifiable time instant is identified from membrane potential traces of the neuron, and guild the synaptic adaptation based on the presynaptic neurons’ contribution at this time instant. The experimental results demonstrate that the proposed algorithm enables the neurons to generate the desired number of spikes, and to detect useful clues embedded within unrelated spiking activities and background noise with a better learning efficiency over the state-of-the-art TDP1 and Multi-Spike Tempotron algorithms. Furthermore, we propose a data-driven dynamic decoding scheme for practical classification tasks, of which the aggregate labels are hard to define. This scheme effectively improves the classification accuracy of the aggregate-label learning algorithms as demonstrated on a speech recognition task. Malu Zhang, Jibin Wu, Yansong Chua, Xiaoling Luo 0001, Zihan Pan, Haizhou Li 0001 |
AAAI | 2 |
| 2019 | Inverted-file R-tree Index Nodes Clustering Based on Multi-objective OptimizationabstractThe evolutionary multi-objective optimization algorithm was used to optimize the clustering and splitting of nodes in the construction of inverted spatial objects index tree(ITSR) in this paper. Considering the objective factors including object's coverage, overlap, group center distance, directory rectangle perimeter and the word similarity between tree nodes, etc., a multi-objective optimization model for solving the optimal ITSR construction is established. To solve the model, this paper find a novel method of constructing high efficiency inverted text-spatial R-tree index. The experimental results show that the algorithm supports the multi-objective optimal clustering of ITSR with high efficiency and accuracy . Wubin Ma, Rui Wang 0017, Weichao Wang, Su Deng, Hongbin Huang, Tao Zhang 0033, Jibin Wu |
CEC | 7 |
| 2019 | Neural Population Coding for Effective Temporal ClassificationabstractNeural encoding plays an important role in faithfully describing the temporally rich patterns, whose instances include human speech and environmental sounds. To classify such spatio-temporal patterns with the Spiking Neural Networks (SNNs), how these patterns are encoded has a direct impact on the complexity of the task. In this paper, we study several existing temporal and population coding schemes in speech and audio recognition. We show that, with population neural coding, the encoded patterns are linearly separable using the Support Vector Machine (SVM). We note that the population neural coding effectively project the temporal information into the spatial domain, thus improving linear separability of the patterns. We achieve an accuracy of 95% and 100% on TIDIGITS and RWCP datasets respectively with SVM classifier. We further implement the Tempotron as an SNN-based classifier on the same datasets and achieve similar results. The study suggests that an effective neural coding scheme is just as important as the classifier. Zihan Pan, Jibin Wu, Malu Zhang, Haizhou Li 0001, Yansong Chua |
IJCNN | 2 |
| 2019 | Deep Spiking Neural Network with Spike Count based Learning RuleabstractDeep spiking neural networks (SNNs) support asynchronous event-driven computation, massive parallelism and demonstrate great potential to improve the energy efficiency of its synchronous analog counterpart. However, insufficient attention has been paid to neural encoding when designing SNN learning rules. Remarkably, the temporal credit assignment has been performed on rate-coded spiking inputs, leading to poor learning efficiency. In this paper, we introduce a novel spike-based learning rule for rate-coded deep SNNs, whereby the spike count of each neuron is used as a surrogate for gradient backpropagation. We evaluate the proposed learning rule by training deep spiking multi-layer perceptron (MLP) and spiking convolutional neural network (CNN) on the UCI machine learning and MNIST handwritten digit datasets. We show that the proposed learning rule achieves state-of-the-art accuracies on all benchmark datasets. The proposed learning rule allows introducing latency, spike rate and hardware constraints into the SNN learning, which is superior to the indirect approach in which conventional artificial neural networks are first trained and then converted to SNNs. Hence, it allows direct deployment to the neuromorphic hardware and supports efficient inference. Notably, a test accuracy of 98.40% was achieved on the MNIST dataset in our experiments with only 10 simulation time steps, when the same latency constraint is imposed during training. Jibin Wu, Yansong Chua, Malu Zhang, Qu Yang, Guoqi Li 0002, Haizhou Li 0001 |
IJCNN | 1 |
| 2019 | Competitive STDP-based Feature Representation Learning for Sound Event ClassificationabstractHumans are good at discriminating environmental sounds and associating them with opportunities or dangers. While the deep learning approach to sound event classification (SEC) is achieving human parity, unsolved problems remain, instances include high computational cost, requirement of massive labeled training data, and question of biological plausibility. Motivated by the human auditory system, we propose a biologically plausible SEC system, which integrates the auditory front-end, population coding, competitive spike-timing-dependent plasticity (STDP) based feature representation learning and supervised temporal classification into a unified spiking neural network (SNN) system. The proposed SEC system achieves a classification accuracy on the RWCP database that is on par with other competitive baseline systems. Furthermore, the STDP-based feature representation learning shows low intra-class variability and high inter-class variability in our experiments, which is highly desirable for pattern classification tasks. Jibin Wu, Malu Zhang, Haizhou Li 0001, Yansong Chua |
IJCNN | 1 |
| 2019 | Robust Sound Recognition: A Neuromorphic Approach
Jibin Wu, Zihan Pan, Malu Zhang, Rohan Kumar Das, Yansong Chua, Haizhou Li 0001 |
INTERSPEECH | 1 |
| 2018 | An Event-Based Cochlear Filter Temporal Encoding Scheme for Speech SignalsabstractSpiking Neural Network (SNN), the third generation of neural networks, has been shown to perform well in pattern recognition tasks involving temporal information, such as speech recognition and motion detection. However, most neural networks, including the SNN, for speech recognition rely on short-time frequency analysis, such as the mel-frequency cepstral coefficients (MFCC), for low-level feature extraction. MFCC feature extraction works by analyzing a window of time signal in multiple frequency bands one window at a time, in a synchronous fashion. This is in contrast to the event-based principle of SNN, whereby electrical impulses are emitted and processed in an asynchronous fashion. Just as speech signals arrive at the human's cochlear filterbank concurrently, but spikes encoding the power in each frequency band are emitted asynchronously, we propose an event-based cochlear filter encoding scheme, whereby the power in each frequency band is directly extracted in the time domain and spikes encoded using the latency code are emitted asynchronously to represent the power of each frequency band. This replaces the traditional MFCC frontend used in most speech recognition models, and makes possible an end-to- end event-based SNN implementation for a speech recognition task. The proposed event-based neural encoding is not only biologically plausible, but also outperforms the MFCC as an encoding frontend for an SNN classifier in a speech recognition task, in terms of higher classification accuracy and lower latency. Such an end-to-end SNN model could be implemented on a neuromorphic chip to fully realize the advantages of event-based processing. Zihan Pan, Haizhou Li 0001, Jibin Wu, Yansong Chua |
IJCNN | 3 |
| 2018 | A Biologically Plausible Speech Recognition Framework Based on Spiking Neural NetworksabstractHumans perform remarkably well for speech recognition using sparse and asynchronous events carried by electrical impulses. Motivated by the observations that human brains primarily learn features from environmental stimuli in an unsupervised manner and consume extremely low power for complex cognitive tasks, we propose a biologically plausible speech recognition mechanism using unsupervised self-organizing map (SOM) for feature representation and event-driven spiking neural network (SNN) for spatiotemporal pattern classification. Moreover, we improve the biological realism of the proposed framework by using mel-scaled filter bank as the front-end, so as to mimic the human auditory system. Our experiments on the TIDIGITS dataset achieve speech recognition accuracy surpassing those of other bio-inspired systems. The proposed SOM-SNN framework can be implemented using the artificial silicon cochlear and neuromorphic processor, so as to fully exploit the potential of event-based speech recognition system. Jibin Wu, Yansong Chua, Haizhou Li 0001 |
IJCNN | 1 |