Yansong Chua

dblp:180/0351 · DBLP profile ↗
← Back
21ranked-venue papers
0as first author
10since 2021 · last 2026
0000-0002-3133-843XORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 17 · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 4 since 2021Systems, architecture and hardware · 1Databases, data management, data science and information retrieval · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
3 papers
Deep learning architectures and training · 42% Efficient and distributed learning · 42% Reinforcement learning · 16%
Computer architecture, parallel and distributed computing, and storage systems
2 papers
Emerging computing paradigms · 100%
Databases, data mining, and information retrieval
1 paper
Data stream processing · 100%
Computer graphics and multimedia
1 paper
Visualization and visual analytics · 100%

Topics — the 12 heaviest of 12, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Efficient and distributed learning
energy-efficient learning
0.912025
MSVIT: Improving Spiking Vision Transformer Using Multi-scale Attention Fusion · IJCAI 2025
Machine learning › Deep learning architectures and training › transformer
vision transformer
0.912025
MSVIT: Improving Spiking Vision Transformer Using Multi-scale Attention Fusion · IJCAI 2025
Data stream processing
stream sampling
0.912025
FPCS: Feature Preserving Compensated Sampling of Streaming Time Series Data · IEEE Trans. Vis. Comput. Graph. 2025
Visualization and visual analytics
time series visualization
0.912025
FPCS: Feature Preserving Compensated Sampling of Streaming Time Series Data · IEEE Trans. Vis. Comput. Graph. 2025
Machine learning › Efficient and distributed learning
neuromorphic computing
0.612022
Training Spiking Neural Networks with Local Tandem Learning · NeurIPS 2022
Machine learning › Efficient and distributed learning › edge computing › on-device machine learning
on-chip learning
0.612022
Training Spiking Neural Networks with Local Tandem Learning · NeurIPS 2022
Machine learning › Deep learning architectures and training
spiking neural network
0.612022
Training Spiking Neural Networks with Local Tandem Learning · NeurIPS 2022
Machine learning › Deep learning architectures and training › spiking neural network
spiking neural network training
0.612022
Training Spiking Neural Networks with Local Tandem Learning · NeurIPS 2022
Emerging computing paradigms
neuromorphic computing
0.622022
MPD-AL: An Efficient Membrane Potential Driven Aggregate-Label Learning Algorithm for Spiking Neurons · AAAI 2019
Training Spiking Neural Networks with Local Tandem Learning · NeurIPS 2022
Machine learning › Reinforcement learning › multi-agent reinforcement learning
credit assignment
0.412019
MPD-AL: An Efficient Membrane Potential Driven Aggregate-Label Learning Algorithm for Spiking Neurons · AAAI 2019
Machine learning › Reinforcement learning › multi-agent reinforcement learning › credit assignment
temporal credit assignment
0.412019
MPD-AL: An Efficient Membrane Potential Driven Aggregate-Label Learning Algorithm for Spiking Neurons · AAAI 2019
Emerging computing paradigms › neuromorphic computing
spiking neural network
0.412019
MPD-AL: An Efficient Membrane Potential Driven Aggregate-Label Learning Algorithm for Spiking Neurons · AAAI 2019

Methods — techniques the papers use, named apart from their topics

feature point compensation · 1.7teacher-student learning · 1.1local tandem learning · 1.1knowledge distillation · 1.1vision transformer · 0.9spiking neural network · 0.9multi-scale attention · 0.9membrane potential driven learning · 0.8dynamic decoding · 0.8
YearPublicationVenuePosition
2026 SPASRNN: boosting spiking recurrent neural networks using sparse gradient decent for semantic representation learning
Yansong Chua, Qian Zhang 0035, Yangyang Shu
Neural Comput. Appl.2
2025 MSVIT: Improving Spiking Vision Transformer Using Multi-scale Attention Fusion
abstract
The combination of Spiking Neural Networks (SNNs) with Vision Transformer architectures has attracted significant attention due to the great potential for energy-efficient and high-performance computing paradigms. However, a substantial performance gap still exists between SNN-based and ANN-based transformer architectures. While existing methods propose spiking self-attention mechanisms that are successfully combined with SNNs, the overall architectures proposed by these methods suffer from a bottleneck in effectively extracting features from different image scales. In this paper, we address this issue and propose MSVIT, a novel spike-driven Transformer architecture, which firstly uses multi-scale spiking attention (MSSA) to enrich the capability of spiking attention blocks. We validate our approach across various main data sets. The experimental results indicate that our MSVIT outperforms existing SNN-based models, positioning itself as a state-of-the-art solution among NN-transformer architectures. The codes are available at https://github.com/Nanhu-AI-Lab/MSViT.
Chenlin Zhou, Jibin Wu, Yansong Chua, Yangyang Shu
IJCAI4
2025 FPCS: Feature Preserving Compensated Sampling of Streaming Time Series Data
abstract
Data visualization aids in making data analysis more intuitive and in-depth, with widespread applications in fields such as biology, finance, and medicine. For massive and continuously growing streaming time series data, these data are typically visualized in the form of line charts, but the data transmission puts significant pressure on the network, leading to visualization lag or even failure to render completely. This paper proposes a universal sampling algorithm FPCS, which retains feature points from continuously received streaming time series data, compensates for the frequent fluctuating feature points, and aims to achieve efficient visualization. This algorithm bridges the gap in sampling for streaming time series data. The algorithm has several advantages: (1) It optimizes the sampling results by compensating for fewer feature points, retaining the visualization features of the original data very well, ensuring high-quality sampled data; (2) The execution time is the shortest compared to similar existing algorithms; (3) It has an almost negligible space overhead; (4) The data sampling process does not depend on the overall data; (5) This algorithm can be applied to infinite streaming data and finite static data.
Yansong Chua
IEEE Trans. Vis. Comput. Graph.3
2024 SGDG: Improving Transformer Seq2Seq Models through Span Generation and Denoise Generation
Zhenfei Yang, Beiming Yu, Chenxiao Dou, Qian Zhang 0035, Yansong Chua
DASFAA (2)5
2024 Spiking Structured State Space Model for Monaural Speech Enhancement
abstract
Speech enhancement seeks to extract clean speech from noisy signals. Traditional deep learning methods face two challenges: efficiently using information in long speech sequences and high computational costs. To address these, we introduce the Spiking Structured State Space Model (SpikingS4). This approach merges the energy efficiency of Spiking Neural Networks (SNN) with the long-range sequence modeling capabilities of Structured State Space Models (S4), offering a compelling solution. Evaluation on the DNS Challenge and VoiceBank+Demand Datasets confirms that Spiking-S4 rivals existing Artificial Neural Network (ANN) methods but with fewer computational resources, as evidenced by reduced parameters and Floating Point Operations (FLOPs).
Yansong Chua
ICASSP3
2024 Event-Driven Spiking Learning Algorithm Using Aggregated Labels
abstract
Traditional spiking learning algorithm aims to train neurons to spike at a specific time or on a particular frequency, which requires precise time and frequency labels in the training process. While in reality, usually only aggregated labels of sequential patterns are provided. The aggregate-label (AL) learning is proposed to discover these predictive features in distracting background streams only by aggregated spikes. It has achieved much success recently, but it is still computationally intensive and has limited use in deep networks. To address these issues, we propose an event-driven spiking aggregate learning algorithm (SALA) in this article. Specifically, to reduce the computational complexity, we improve the conventional spike-threshold-surface (STS) calculation in AL learning by analytical calculating voltage peak values in spiking neurons. Then we derive the algorithm to multilayers by event-driven strategy using aggregated spikes. We conduct comprehensive experiments on various tasks including temporal clue recognition, segmented and continuous speech recognition, and neuromorphic image classification. The experimental results demonstrate that the new STS method improves the efficiency of AL learning significantly, and the proposed algorithm outperforms the conventional spiking algorithm in various temporal clue recognition tasks.
Xiurui Xie, Yansong Chua, Guisong Liu, Malu Zhang, Guangchun Luo, Huajin Tang
IEEE Trans. Neural Networks Learn. Syst.2
2023 Adaptive Axonal Delays in Feedforward Spiking Neural Networks for Accurate Spoken Word Recognition
abstract
Spiking neural networks (SNN) are a promising research avenue for building accurate and efficient automatic speech recognition systems. Recent advances in audio-to-spike encoding and training algorithms enable SNN to be applied in practical tasks. Biologically-inspired SNN communicates using sparse asynchronous events. Therefore, spike-timing is critical to SNN performance. In this aspect, most works focus on training synaptic weights and few have considered delays in event transmission, namely axonal delay. In this work, we consider a learnable axonal delay capped at a maximum value, which can be adapted according to the axonal delay distribution in each network layer. We show that our proposed method achieves the best classification results reported on the SHD dataset (92.45%) and NTIDIGITS dataset (95.09%). Our work illustrates the potential of training axonal delays for tasks with complex temporal structures.
Pengfei Sun 0003, Ehsan Eqlimi, Yansong Chua, Paul Devos, Dick Botteldooren
ICASSP3
2023 A Tandem Learning Rule for Effective Training and Rapid Inference of Deep Spiking Neural Networks
abstract
Spiking neural networks (SNNs) represent the most prominent biologically inspired computing model for neuromorphic computing (NC) architectures. However, due to the nondifferentiable nature of spiking neuronal functions, the standard error backpropagation algorithm is not directly applicable to SNNs. In this work, we propose a tandem learning framework that consists of an SNN and an artificial neural network (ANN) coupled through weight sharing. The ANN is an auxiliary structure that facilitates the error backpropagation for the training of the SNN at the spike-train level. To this end, we consider the spike count as the discrete neural representation in the SNN and design an ANN neuronal activation function that can effectively approximate the spike count of the coupled SNN. The proposed tandem learning rule demonstrates competitive pattern recognition and regression capabilities on both the conventional frame- and event-based vision datasets, with at least an order of magnitude reduced inference time and total synaptic operations over other state-of-the-art SNN implementations. Therefore, the proposed tandem learning rule offers a novel solution to training efficient, low latency, and high-accuracy deep SNNs with low computing resources.
Jibin Wu, Yansong Chua, Malu Zhang, Guoqi Li 0002, Haizhou Li 0001, Kay Chen Tan
IEEE Trans. Neural Networks Learn. Syst.2
2022 Training Spiking Neural Networks with Local Tandem Learning
abstract
Spiking neural networks (SNNs) are shown to be more biologically plausible and energy efficient over their predecessors. However, there is a lack of an efficient and generalized training method for deep SNNs, especially for deployment on analog computing substrates. In this paper, we put forward a generalized learning rule, termed Local Tandem Learning (LTL). The LTL rule follows the teacher-student learning approach by mimicking the intermediate feature representations of a pre-trained ANN. By decoupling the learning of network layers and leveraging highly informative supervisor signals, we demonstrate rapid network convergence within five training epochs on the CIFAR-10 dataset while having low computational complexity. Our experimental results have also shown that the SNNs thus trained can achieve comparable accuracies to their teacher ANNs on CIFAR-10, CIFAR-100, and Tiny ImageNet datasets. Moreover, the proposed LTL rule is hardware friendly. It can be easily implemented on-chip to perform fast parameter calibration and provide robustness against the notorious device non-ideality issues. It, therefore, opens up a myriad of opportunities for training and deployment of SNN on ultra-low-power mixed-signal neuromorphic computing chips.
Qu Yang, Jibin Wu, Malu Zhang, Yansong Chua, Xinchao Wang, Haizhou Li 0001
NeurIPS4
2022 Rectified Linear Postsynaptic Potential Function for Backpropagation in Deep Spiking Neural Networks
abstract
Spiking neural networks (SNNs) use spatiotemporal spike patterns to represent and transmit information, which are not only biologically realistic but also suitable for ultralow-power event-driven neuromorphic implementation. Just like other deep learning techniques, deep SNNs (DeepSNNs) benefit from the deep architecture. However, the training of DeepSNNs is not straightforward because the well-studied error backpropagation (BP) algorithm is not directly applicable. In this article, we first establish an understanding as to why error BP does not work well in DeepSNNs. We then propose a simple yet efficient rectified linear postsynaptic potential function (ReL-PSP) for spiking neurons and a spike-timing-dependent BP (STDBP) learning algorithm for DeepSNNs where the timing of individual spikes is used to convey information (temporal coding), and learning (BP) is performed based on spike timing in an event-driven manner. We show that DeepSNNs trained with the proposed single spike time-based learning algorithm can achieve the state-of-the-art classification accuracy. Furthermore, by utilizing the trained model parameters obtained from the proposed STDBP learning algorithm, we demonstrate ultralow-power inference operations on a recently proposed neuromorphic inference accelerator. The experimental results also show that the neuromorphic hardware consumes 0.751 mW of the total power consumption and achieves a low latency of 47.71 ms to classify an image from the Modified National Institute of Standards and Technology (MNIST) dataset. Overall, this work investigates the contribution of spike timing dynamics for information encoding, synaptic plasticity, and decision-making, providing a new perspective to the design of future DeepSNNs and neuromorphic hardware.
Malu Zhang, Jibin Wu, Ammar Belatreche, Burin Amornpaisannon, Venkata Pavan Kumar Miriyala, Hong Qu 0002, Yansong Chua, Trevor E. Carlson, Haizhou Li 0001
IEEE Trans. Neural Networks Learn. Syst.9
2020 Classifying Neuromorphic Datasets with Tempotron and Spike Timing Dependent Plasticity
abstract
Although the spike rate of a neuron codes useful information, there is a lot of evidence that information is contained in the precise timing of spikes. Static images have long been used as benchmarks for ANNs. However, in the neuromorphic community novel benchmarks have developed that have both spatial and temporal information. We require capable SNN algorithms that classify temporal datasets. There has been a lot of research in training SNNs. Trained ANNs have been converted to SNNs. SNNs have been trained using variants of backpropagation. However, less research has been done on local learning rules, and biologically plausible methods such as spike timing dependent plasticity (STDP) in training SNNs. In this paper, we present a biologically plausible algorithm that utilizes STDP and tempotron learning rule to classify temporal datasets. We have used DvsGesture to test our algorithm.
Laxmi R. Iyer, Yansong Chua
IJCNN2
2020 Fast Texture Classification Using Tactile Neural Coding and Spiking Neural Network
abstract
Touch is arguably the most important sensing modality in physical interactions. However, tactile sensing has been largely under-explored in robotics applications owing to the complexity in making perceptual inferences until the recent advancements in machine learning or deep learning in particular. Touch perception is strongly influenced by both its temporal dimension similar to audition and its spatial dimension similar to vision. While spatial cues can be learned episodically, temporal cues compete against the system's re-sponse/reaction time to provide accurate inferences. In this paper, we propose a fast tactile-based texture classification framework which makes use of the spiking neural network to learn from the neural coding of the conventional tactile sensor readings. The framework is implemented and tested on two independent tactile datasets collected in sliding motion on 20 material textures. Our results show that the framework is able to make much more accurate inferences ahead of time as compared to that by the state-of-the-art learning approaches.
Tasbolat Taunyazov, Yansong Chua, Ruihan Gao, Harold Soh, Yan Wu 0002
IROS2
2020 Supervised learning in spiking neural networks with synaptic delay-weight plasticity
Malu Zhang, Jibin Wu, Ammar Belatreche, Zihan Pan, Xiurui Xie, Yansong Chua, Guoqi Li 0002, Hong Qu 0002, Haizhou Li 0001
Neurocomputing6
2019 MPD-AL: An Efficient Membrane Potential Driven Aggregate-Label Learning Algorithm for Spiking Neurons
abstract
One of the long-standing questions in biology and machine learning is how neural networks may learn important features from the input activities with a delayed feedback, commonly known as the temporal credit-assignment problem. The aggregate-label learning is proposed to resolve this problem by matching the spike count of a neuron with the magnitude of a feedback signal. However, the existing threshold-driven aggregate-label learning algorithms are computationally intensive, resulting in relatively low learning efficiency hence limiting their usability in practical applications. In order to address these limitations, we propose a novel membrane-potential driven aggregate-label learning algorithm, namely MPD-AL. With this algorithm, the easiest modifiable time instant is identified from membrane potential traces of the neuron, and guild the synaptic adaptation based on the presynaptic neurons’ contribution at this time instant. The experimental results demonstrate that the proposed algorithm enables the neurons to generate the desired number of spikes, and to detect useful clues embedded within unrelated spiking activities and background noise with a better learning efficiency over the state-of-the-art TDP1 and Multi-Spike Tempotron algorithms. Furthermore, we propose a data-driven dynamic decoding scheme for practical classification tasks, of which the aggregate labels are hard to define. This scheme effectively improves the classification accuracy of the aggregate-label learning algorithms as demonstrated on a speech recognition task.
Malu Zhang, Jibin Wu, Yansong Chua, Xiaoling Luo 0001, Zihan Pan, Haizhou Li 0001
AAAI3
2019 Neural Population Coding for Effective Temporal Classification
abstract
Neural encoding plays an important role in faithfully describing the temporally rich patterns, whose instances include human speech and environmental sounds. To classify such spatio-temporal patterns with the Spiking Neural Networks (SNNs), how these patterns are encoded has a direct impact on the complexity of the task. In this paper, we study several existing temporal and population coding schemes in speech and audio recognition. We show that, with population neural coding, the encoded patterns are linearly separable using the Support Vector Machine (SVM). We note that the population neural coding effectively project the temporal information into the spatial domain, thus improving linear separability of the patterns. We achieve an accuracy of 95% and 100% on TIDIGITS and RWCP datasets respectively with SVM classifier. We further implement the Tempotron as an SNN-based classifier on the same datasets and achieve similar results. The study suggests that an effective neural coding scheme is just as important as the classifier.
Zihan Pan, Jibin Wu, Malu Zhang, Haizhou Li 0001, Yansong Chua
IJCNN5
2019 Deep Spiking Neural Network with Spike Count based Learning Rule
abstract
Deep spiking neural networks (SNNs) support asynchronous event-driven computation, massive parallelism and demonstrate great potential to improve the energy efficiency of its synchronous analog counterpart. However, insufficient attention has been paid to neural encoding when designing SNN learning rules. Remarkably, the temporal credit assignment has been performed on rate-coded spiking inputs, leading to poor learning efficiency. In this paper, we introduce a novel spike-based learning rule for rate-coded deep SNNs, whereby the spike count of each neuron is used as a surrogate for gradient backpropagation. We evaluate the proposed learning rule by training deep spiking multi-layer perceptron (MLP) and spiking convolutional neural network (CNN) on the UCI machine learning and MNIST handwritten digit datasets. We show that the proposed learning rule achieves state-of-the-art accuracies on all benchmark datasets. The proposed learning rule allows introducing latency, spike rate and hardware constraints into the SNN learning, which is superior to the indirect approach in which conventional artificial neural networks are first trained and then converted to SNNs. Hence, it allows direct deployment to the neuromorphic hardware and supports efficient inference. Notably, a test accuracy of 98.40% was achieved on the MNIST dataset in our experiments with only 10 simulation time steps, when the same latency constraint is imposed during training.
Jibin Wu, Yansong Chua, Malu Zhang, Qu Yang, Guoqi Li 0002, Haizhou Li 0001
IJCNN2
2019 Competitive STDP-based Feature Representation Learning for Sound Event Classification
abstract
Humans are good at discriminating environmental sounds and associating them with opportunities or dangers. While the deep learning approach to sound event classification (SEC) is achieving human parity, unsolved problems remain, instances include high computational cost, requirement of massive labeled training data, and question of biological plausibility. Motivated by the human auditory system, we propose a biologically plausible SEC system, which integrates the auditory front-end, population coding, competitive spike-timing-dependent plasticity (STDP) based feature representation learning and supervised temporal classification into a unified spiking neural network (SNN) system. The proposed SEC system achieves a classification accuracy on the RWCP database that is on par with other competitive baseline systems. Furthermore, the STDP-based feature representation learning shows low intra-class variability and high inter-class variability in our experiments, which is highly desirable for pattern classification tasks.
Jibin Wu, Malu Zhang, Haizhou Li 0001, Yansong Chua
IJCNN4
2019 Robust Sound Recognition: A Neuromorphic Approach
Jibin Wu, Zihan Pan, Malu Zhang, Rohan Kumar Das, Yansong Chua, Haizhou Li 0001
INTERSPEECH5
2018 Classifying Neuromorphic Data Using a Deep Learning Framework for Image Classification
abstract
In the field of artificial intelligence, neuromorphic computing has been around for several decades. Deep learning has however made much recent progress such that it consistently outperforms neuromorphic learning algorithms in classification tasks in terms of accuracy. Specifically in the field of image classification, neuromorphic computing has been traditionally using either the temporal or rate code for encoding static images in datasets into spike trains. It is only till recently, that neuromorphic vision sensors are widely used by the neuromorphic research community, and provides an alternative to such encoding methods. Since then, several neuromorphic datasets are obtained by applying such sensors on image datasets (e.g. the neuromorphic CALTECH 101) have been introduced. These data are encoded in spike trains and hence seem ideal for benchmarking of neuromorphic learning algorithms. Specifically, we train a deep learning framework used for image classification on the CALTECH 101 and a collapsed version of the neuromorphic CALTECH 101 datasets. We obtained an accuracy of 91.66% and 78.01% for the CALTECH 101 and neuromorphic CALTECH 101 datasets respectively. For CALTECH 101, our accuracy is close to the best reported accuracy, while for neuromorphic CALTECH 101, it outperforms the last best reported accuracy by over 10%. This raises the question of the suitability of such datasets as benchmarks for neuromorphic learning algorithms.
Roshan Gopalakrishnan, Yansong Chua, Laxmi R. Iyer
ICARCV2
2018 An Event-Based Cochlear Filter Temporal Encoding Scheme for Speech Signals
abstract
Spiking Neural Network (SNN), the third generation of neural networks, has been shown to perform well in pattern recognition tasks involving temporal information, such as speech recognition and motion detection. However, most neural networks, including the SNN, for speech recognition rely on short-time frequency analysis, such as the mel-frequency cepstral coefficients (MFCC), for low-level feature extraction. MFCC feature extraction works by analyzing a window of time signal in multiple frequency bands one window at a time, in a synchronous fashion. This is in contrast to the event-based principle of SNN, whereby electrical impulses are emitted and processed in an asynchronous fashion. Just as speech signals arrive at the human's cochlear filterbank concurrently, but spikes encoding the power in each frequency band are emitted asynchronously, we propose an event-based cochlear filter encoding scheme, whereby the power in each frequency band is directly extracted in the time domain and spikes encoded using the latency code are emitted asynchronously to represent the power of each frequency band. This replaces the traditional MFCC frontend used in most speech recognition models, and makes possible an end-to- end event-based SNN implementation for a speech recognition task. The proposed event-based neural encoding is not only biologically plausible, but also outperforms the MFCC as an encoding frontend for an SNN classifier in a speech recognition task, in terms of higher classification accuracy and lower latency. Such an end-to-end SNN model could be implemented on a neuromorphic chip to fully realize the advantages of event-based processing.
Zihan Pan, Haizhou Li 0001, Jibin Wu, Yansong Chua
IJCNN4
2018 A Biologically Plausible Speech Recognition Framework Based on Spiking Neural Networks
abstract
Humans perform remarkably well for speech recognition using sparse and asynchronous events carried by electrical impulses. Motivated by the observations that human brains primarily learn features from environmental stimuli in an unsupervised manner and consume extremely low power for complex cognitive tasks, we propose a biologically plausible speech recognition mechanism using unsupervised self-organizing map (SOM) for feature representation and event-driven spiking neural network (SNN) for spatiotemporal pattern classification. Moreover, we improve the biological realism of the proposed framework by using mel-scaled filter bank as the front-end, so as to mimic the human auditory system. Our experiments on the TIDIGITS dataset achieve speech recognition accuracy surpassing those of other bio-inspired systems. The proposed SOM-SNN framework can be implemented using the artificial silicon cochlear and neuromorphic processor, so as to fully exploit the potential of event-based speech recognition system.
Jibin Wu, Yansong Chua, Haizhou Li 0001
IJCNN2