Zhengyu Ma

dblp:242/2942 · DBLP profile ↗
← Back
26ranked-venue papers
2as first author
26since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 18 · 1 first-author · 18 since 2021Graphics, computer vision, multimedia, augmented reality and games · 11 · 1 first-author · 11 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Parallel Training Time-to-First-Spike Spiking Neural Networks
abstract
Spiking Neural Networks (SNNs) offer a promising energy-efficient computing paradigm owing to their event-driven properties and biologically inspired dynamics. Among various encoding schemes, Time-to-First-Spike (TTFS) is particularly notable for its extreme sparsity, utilizing a single spike per neuron to maximize energy efficiency. However, two significant challenges persist: effectively leveraging TTFS sparsity to minimize training costs on Graphics Processing Units (GPUs), and bridging the performance gap between TTFS-based SNNs and their rate-based counterparts. To address these issues, we propose a parallel training algorithm for accelerated execution and a novel decoding strategy for enhanced performance. Specifically, we derive both forward and backward propagation equations for parallelized TTFS SNNs, enabling precise calculation of first-spike timings and gradients. Furthermore, we analyze the limitations of existing output decoders and introduce a membrane potential–based decoder, complemented by an incremental time-step training strategy, to improve accuracy. Our approach achieves state-of-the-art accuracy for TTFS SNNs on several benchmarks, including MNIST (99.51%), Fashion-MNIST (93.14%), CIFAR-10 (95.06%), and CIFAR-100 (74.07%).
Kaiwei Che, Wei Fang 0006, Yifan Huang 0002, Zhengyu Ma, Yonghong Tian 0001
AAAI5
2026 SpikCommander: A High-performance Spiking Transformer with Multi-view Learning for Efficient Speech Command Recognition
abstract
Spiking neural networks (SNNs) offer a promising path toward energy-efficient speech command recognition (SCR) by leveraging their event-driven processing paradigm. However, existing SNN-based SCR methods often struggle to capture rich temporal dependencies and contextual information from speech due to limited temporal modeling and binary spike-based representations. To address these challenges, we first introduce the multi-view spiking temporal-aware self-attention (MSTASA) module, which combines effective spiking temporal-aware attention with a multi-view learning framework to model complementary temporal dependencies in speech commands. Building on MSTASA, we further propose SpikCommander, a fully spike-driven transformer architecture that integrates MSTASA with a spiking contextual refinement channel MLP (SCR-MLP) to jointly enhance temporal context modeling and channel-wise feature integration. We evaluate our method on three benchmark datasets: the Spiking Heidelberg Dataset (SHD), the Spiking Speech Commands (SSC), and the Google Speech Commands V2 (GSC). Extensive experiments demonstrate that SpikCommander consistently outperforms state-of-the-art (SOTA) SNN approaches with fewer parameters under comparable time steps, highlighting its effectiveness and efficiency for robust speech command recognition.
Jiaqi Wang 0003, Liutao Yu, Xiongri Shen, Sihang Guo, Chenlin Zhou, Zhiguo Zhang 0001, Zhengyu Ma
AAAI9
2026 Spikingformer: A Key Foundation Model for Spiking Neural Networks
abstract
Spiking neural networks (SNNs) offer a promising energy-efficient alternative to artificial neural networks, due to their event-driven spiking computation. However, some foundation SNN backbones (including Spikformer and SEW ResNet) suffer from non-spike computations (integer-float multiplications) caused by the structure of their residual connections. These non-spike computations increase SNNs' power consumption and make them unsuitable for deployment on mainstream neuromorphic hardware. In this paper, we analyze the spike-driven behavior of the residual connection methods in SNNs. We then present Spikingformer, a novel spiking transformer backbone that merges the MS Residual connection with Self-Attention in a biologically plausible way to address the non-spike computation challenge in Spikformer while maintaining global modeling capabilities. We evaluate Spikingformer across 13 datasets spanning large static images, neuromorphic data, and natural language tasks, and demonstrate the effectiveness and universality of Spikingformer, setting a vital benchmark for spiking neural networks. In addition, with the spike-driven features and global modeling capabilities, Spikingformer is expected to become a more efficient general-purpose SNN backbone towards energy-efficient artificial intelligence.
Chenlin Zhou, Liutao Yu, Zhaokun Zhou, Han Zhang 0035, Jiaqi Wang 0003, Zhengyu Ma, Yonghong Tian 0001
AAAI7
2026 Improving Pseudo-Labeling and Representation Balance in Realistic Long-Tailed Semi-Supervised Learning
abstract
Despite the remarkable progress of semi-supervised learning (SSL), its effectiveness under realistic long-tailed settings remains limited. In such settings, labeled data is severely imbalanced, while the distribution of unlabeled data is unknown and often mismatched. Under these conditions, class imbalance inherently leads to biased decision boundaries during training, and distribution mismatch causes unreliable pseudo-labels that further exacerbate this bias. Moreover, realistic long-tailed semi-supervised learning suffers from representation imbalance in feature learning, where dominant classes occupy large regions of the feature space while minority classes become overly compact. To address these challenges, we propose PRB-SSL, a method for improving pseudo-label reliability and representation balance in realistic long-tailed semi-supervised learning. PRB-SSL is built upon a dual-branch framework. Specifically, a Biased Predictor adapts to the unlabeled data distribution to generate more reliable pseudo-labels under distribution mismatch, while a Balanced Predictor with decision-level rebalancing mitigates class-imbalance-induced boundary bias and enables balanced inference. Furthermore, PRB-SSL introduces learning status, a dynamic class-level measure, to regulate feature diffusion during semi-supervised learning, suppressing excessive expansion of well-learned classes while preserving exploration for under-learned ones, thereby alleviating representation imbalance. Extensive experiments on CIFAR-10-LT, CIFAR-100-LT, and STL-10-LT demonstrate that PRB-SSL consistently outperforms state-of-the-art methods under realistic long-tailed semi-supervised learning settings.
Zhengyu Ma, Qifeng Zhou
ICMR1
2026 The limits of bio-molecular modeling with large language models: a cross-scale evaluation
abstract
MOTIVATION: The modeling of bio-molecular system across molecular scales remains a central challenge in scientific research. Large language models (LLMs) are increasingly applied to bio-molecular discovery, yet systematic evaluation across multi-scale biological problems and rigorous assessment of their tool-augmented capabilities remain limited. RESULTS: We reveal a systematic gap between LLM performance and mechanistic understanding through the proposed cross-scale bio-molecular benchmark: BioMol-LLM-Bench, a unified framework comprising 26 downstream tasks that covers 4 distinct difficulty levels, and computational tools are integrated for a more comprehensive evaluation. Evaluation on 13 representative models reveals 4 benchmark-specific observations: chain-of-thought-style training does not consistently improve performance on the evaluated biological tasks; the evaluated hybrid mamba-attention model shows strong performance on long bio-molecular sequence tasks; supervised fine-tuned models show task-specific specialization with reduced performance in some general settings; and current LLMs perform better on classification tasks than on challenging regression tasks under this benchmark setting. AVAILABILITY: Source code is available at https://github.com/AI-HPC-Research-Team/BioMol-LLM-Bench.
Yaxin Xu, Yue Zhou 0014, Zhengyu Ma, Fengwei An, Zhixiang Ren
Bioinform.4
2026 Efficient speech command recognition leveraging spiking neural networks and progressive time-scaled curriculum distillation
Jiaqi Wang 0003, Liutao Yu, Liwei Huang, Chenlin Zhou, Han Zhang 0035, Zhenxi Song, Honghai Liu 0001, Min Zhang 0005, Zhengyu Ma, Zhiguo Zhang 0001
Neural Networks9
2025 Enhancing GraphRAG with Syntactic Knowledge and Mixture-of-Experts for Knowledge-Intensive QA
abstract
Knowledge-intensive question answering (QA) requires deep reasoning over heterogeneous sources and factually consistent answers, but existing RAG and GraphRAG frameworks are limited by surface-level semantic similarity-causing multi-hop inference and entity disambiguation errors. This work proposes a GraphRAG enhancement integrating two core components: 1) syntactic knowledge infusion (via dependency parsing and structural prompts) to align retrieval/context evidence with complex questions' compositional logic; 2) a Mixture-of-Experts (MoE) module for dynamic expert selection, optimizing multi-hop reasoning and cross-domain generalization. Evaluations on QALD-h, DBpedia QA, and MetaQA show state-of-the-art performance across exact match, factual consistency, and structural fidelity. Ablation/qualitative analyses confirm that syntactic scaffolding plus expert specialization boosts accuracy, interpretability, robustness, and generalizability. This study establishes a new knowledge-intensive QA paradigm, highlighting the value of unifying explicit structural modeling with dynamic modular reasoning.
Zhengyu Ma, Xiaojia Jin, Bo Wang 0011
ICTAI1
2025 DSF-Net: Dynamic Sparse Fusion of Event-RGB via Spike-Triggered Attention for High-Speed Detection
Dongyang Ma, Zhengyu Ma, Wei Zhang 0161, Yonghong Tian 0001
ACM Multimedia2
2025 Time-Evolving Dynamical System for Learning Latent Representations of Mouse Visual Neural Activity
abstract
Seeking high-quality representations with latent variable models (LVMs) to reveal the intrinsic correlation between neural activity and behavior or sensory stimuli has attracted much interest. In the study of the biological visual system, naturalistic visual stimuli are inherently high-dimensional and time-dependent, leading to intricate dynamics within visual neural activity. However, most work on LVMs has not explicitly considered neural temporal relationships. To cope with such conditions, we propose Time-Evolving Visual Dynamical System (TE-ViDS), a sequential LVM that decomposes neural activity into low-dimensional latent representations that evolve over time. To better align the model with the characteristics of visual neural activity, we split latent representations into two parts and apply contrastive learning to shape them. Extensive experiments on synthetic datasets and real neural datasets from the mouse visual cortex demonstrate that TE-ViDS achieves the best decoding performance on naturalistic scenes/movies, extracts interpretable latent trajectories that uncover clear underlying neural dynamics, and provides new insights into differences in visual information processing between subjects and between cortical regions. In summary, TE-ViDS is markedly competent in extracting stimulus-relevant embeddings from visual neural activity and contributes to the understanding of visual processing mechanisms. Our codes are available at https://github.com/Grasshlw/Time-Evolving-Visual-Dynamical-System.
Liwei Huang, Zhengyu Ma, Liutao Yu, Yonghong Tian 0001
NeurIPS2
2025 S$^2$M-Former: Spiking Symmetric Mixing Branchformer for Brain Auditory Attention Detection
abstract
Auditory attention detection (AAD) aims to decode listeners' focus in complex auditory environments from electroencephalography (EEG) recordings, which is crucial for developing neuro-steered hearing devices. Despite recent advancements, EEG-based AAD remains hindered by the absence of synergistic frameworks that can fully leverage complementary EEG features under energy-efficiency constraints. We propose ***S$^2$M-Former***, a novel ***s***piking ***s***ymmetric ***m***ixing framework to address this limitation through two key innovations: i) Presenting a spike-driven symmetric architecture composed of parallel spatial and frequency branches with mirrored modular design, leveraging biologically plausible token-channel mixers to enhance complementary learning across branches; ii) Introducing lightweight 1D token sequences to replace conventional 3D operations, reducing parameters by 14.7$\times$. The brain-inspired spiking architecture further reduces power consumption, achieving a 5.8$\times$ energy reduction compared to recent ANN methods, while also surpassing existing SNN baselines in terms of parameter efficiency and performance. Comprehensive experiments on three AAD benchmarks (KUL, DTU and AV-GC-AAD) across three settings (within-trial, cross-trial and cross-subject) demonstrate that S$^2$M-Former achieves comparable state-of-the-art (SOTA) decoding accuracy, making it a promising low-power, high-performance solution for AAD tasks. Code is available at https://github.com/JackieWang9811/S2M-Former.
Jiaqi Wang 0003, Zhengyu Ma, Xiongri Shen, Chenlin Zhou, Han Zhang 0035, Zhenxi Song, Zhiguo Zhang 0001
NeurIPS2
2025 Multiplication-Free Parallelizable Spiking Neurons with Efficient Spatio-Temporal Dynamics
abstract
Spiking Neural Networks (SNNs) are distinguished from Artificial Neural Networks (ANNs) for their complex neuronal dynamics and sparse binary activations (spikes) inspired by the biological neural system. Traditional neuron models use iterative step-by-step dynamics, resulting in serial computation and slow training speed of SNNs. Recently, parallelizable spiking neuron models have been proposed to fully utilize the massive parallel computing ability of graphics processing units to accelerate the training of SNNs. However, existing parallelizable spiking neuron models involve dense floating operations and can only achieve high long-term dependencies learning ability with a large order at the cost of huge computational and memory costs. To solve the dilemma of performance and costs, we propose the mul-free channel-wise Parallel Spiking Neuron, which is hardware-friendly and suitable for SNNs’ resource-restricted application scenarios. The proposed neuron imports the channel-wise convolution to enhance the learning ability, induces the sawtooth dilations to reduce the neuron order, and employs the bit-shift operation to avoid multiplications. The algorithm for the design and implementation of acceleration methods is discussed extensively. Our methods are validated in neuromorphic Spiking Heidelberg Digits voices, sequential CIFAR images, and neuromorphic DVS-Lip vision datasets, achieving superior performance over SOTA spiking neurons. Training speed results demonstrate the effectiveness of our acceleration methods, providing a practical reference for future research. Our code is available at Github.
Wei Fang 0006, Zhengyu Ma, Zihan Huang, Zhaokun Zhou, Yonghong Tian 0001, Timothée Masquelier
NeurIPS3
2025 High-Rate Monocular Depth Estimation via Cross Frame-Rate Collaboration of Frames and Events
Xu Liu 0006, Xiaopeng Fan 0001, Jianing Li 0001, Dianze Li, Wei Zhang 0161, Zhengyu Ma, Yonghong Tian 0001
Int. J. Comput. Vis.6
2024 Enhancing EEG-to-Text Decoding through Transferable Representations from Pre-trained Contrastive EEG-Text Masked Autoencoder
abstract
Reconstructing natural language from noninvasive electroencephalography (EEG) holds great promise as a language decoding technology for brain-computer interfaces (BCIs).How-
Jiaqi Wang 0003, Zhenxi Song, Zhengyu Ma, Xipeng Qiu, Min Zhang 0005, Zhiguo Zhang 0001
ACL (1)3
2024 Temporal Contrastive Learning for Spiking Neural Networks
Haonan Qiu, Zeyin Song, Yanqi Chen, Munan Ning, Wei Fang 0006, Zhengyu Ma, Li Yuan 0007, Yonghong Tian 0001
ICANN (10)7
2024 Long-Range Feedback Spiking Network Captures Dynamic and Static Representations of the Visual Cortex under Movie Stimuli
abstract
Deep neural networks (DNNs) are widely used models for investigating biological visual representations. However, existing DNNs are mostly designed to analyze neural responses to static images, relying on feedforward structures and lacking physiological neuronal mechanisms. There is limited insight into how the visual cortex represents natural movie stimuli that contain context-rich information. To address these problems, this work proposes the long-range feedback spiking network (LoRaFB-SNet), which mimics top-down connections between cortical regions and incorporates spike information processing mechanisms inherent to biological neurons. Taking into account the temporal dependence of representations under movie stimuli, we present Time-Series Representational Similarity Analysis (TSRSA) to measure the similarity between model representations and visual cortical representations of mice. LoRaFB-SNet exhibits the highest level of representational similarity, outperforming other well-known and leading alternatives across various experimental paradigms, especially when representing long movie stimuli. We further conduct experiments to quantify how temporal structures (dynamic information) and static textures (static information) of the movie stimuli influence representational similarity, suggesting that our model benefits from long-range feedback to encode context-dependent representations just like the brain. Altogether, LoRaFB-SNet is highly competent in capturing both dynamic and static representations of the mouse visual cortex and contributes to the understanding of movie processing mechanisms of the visual system. Our codes are available at https://github.com/Grasshlw/SNN-Neural-Similarity-Movie.
Liwei Huang, Zhengyu Ma, Liutao Yu, Yonghong Tian 0001
NeurIPS2
2024 QKFormer: Hierarchical Spiking Transformer using Q-K Attention
abstract
Spiking Transformers, which integrate Spiking Neural Networks (SNNs) with Transformer architectures, have attracted significant attention due to their potential for low energy consumption and high performance. However, there remains a substantial gap in performance between SNNs and Artificial Neural Networks (ANNs). To narrow this gap, we have developed QKFormer, a direct training spiking transformer with the following features: i) _Linear complexity and high energy efficiency_, the novel spike-form Q-K attention module efficiently models the token or channel attention through binary vectors and enables the construction of larger models. ii) _Multi-scale spiking representation_, achieved by a hierarchical structure with the different numbers of tokens across blocks. iii) _Spiking Patch Embedding with Deformed Shortcut (SPEDS)_, enhances spiking information transmission and integration, thus improving overall performance. It is shown that QKFormer achieves significantly superior performance over existing state-of-the-art SNN models on various mainstream datasets. Notably, with comparable size to Spikformer (66.34 M, 74.81\%), QKFormer (64.96 M) achieves a groundbreaking top-1 accuracy of **85.65\%** on ImageNet-1k, substantially outperforming Spikformer by **10.84\%**. To our best knowledge, this is the first time that directly training SNNs have exceeded 85\% accuracy on ImageNet-1K.
Chenlin Zhou, Han Zhang 0035, Zhaokun Zhou, Liutao Yu, Liwei Huang, Xiaopeng Fan 0001, Li Yuan 0007, Zhengyu Ma, Yonghong Tian 0001
NeurIPS8
2024 Self-architectural knowledge distillation for spiking neural networks
Haonan Qiu, Munan Ning, Zeyin Song, Wei Fang 0006, Yanqi Chen, Zhengyu Ma, Li Yuan 0007, Yonghong Tian 0001
Neural Networks7
2023 Deep Spiking Neural Networks with High Representation Similarity Model Visual Pathways of Macaque and Mouse
abstract
Deep artificial neural networks (ANNs) play a major role in modeling the visual pathways of primate and rodent. However, they highly simplify the computational properties of neurons compared to their biological counterparts. Instead, Spiking Neural Networks (SNNs) are more biologically plausible models since spiking neurons encode information with time sequences of spikes, just like biological neurons do. However, there is a lack of studies on visual pathways with deep SNNs models. In this study, we model the visual cortex with deep SNNs for the first time, and also with a wide range of state-of-the-art deep CNNs and ViTs for comparison. Using three similarity metrics, we conduct neural representation similarity experiments on three neural datasets collected from two species under three types of stimuli. Based on extensive similarity analyses, we further investigate the functional hierarchy and mechanisms across species. Almost all similarity scores of SNNs are higher than their counterparts of CNNs with an average of 6.6%. Depths of the layers with the highest similarity scores exhibit little differences across mouse cortical regions, but vary significantly across macaque regions, suggesting that the visual processing structure of mice is more regionally homogeneous than that of macaques. Besides, the multi-branch structures observed in some top mouse brain-like neural networks provide computational evidence of parallel processing streams in mice, and the different performance in fitting macaque neural representations under different stimuli exhibits the functional specialization of information processing in macaques. Taken together, our study demonstrates that SNNs could serve as promising candidates to better model and explain the functional hierarchy and mechanisms of the visual system.
Liwei Huang, Zhengyu Ma, Liutao Yu, Yonghong Tian 0001
AAAI2
2023 A Unified Framework for Soft Threshold Pruning
Yanqi Chen, Zhengyu Ma, Wei Fang 0006, Xiawu Zheng, Zhaofei Yu, Yonghong Tian 0001
ICLR2
2023 Modality-Fusion Spiking Transformer Network for Audio-Visual Zero-Shot Learning
abstract
Audio-visual zero-shot learning (ZSL), which learns to classify video data from the classes not being observed during training, is challenging. In audio-visual ZSL, both semantic and temporal information from different modalities is relevant to each other. However, effectively extracting and fusing information from audio and visual remains an open challenge. In this work, we propose an Audio-Visual Modality-fusion Spiking Transformer network (AVMST) for audio-visual ZSL. To be more specific, AVMST provides a spiking neural network (SNN) module for extracting conspicuous temporal information of each modality, a cross-attention block to effectively fuse the temporal and semantic information, and a transformer reasoning module to further explore the interrelationships of fusion features. To provide robust temporal features, the spiking threshold of the SNN module is adjusted dynamically based on the semantic cues of different modalities. The generated feature map is in accordance with the zero-shot learning property thanks to our proposed spiking transformer’s ability to combine the robustness of SNN feature extraction and the precision of transformer feature inference. Extensive experiments on three benchmark audio-visual datasets (i.e., VGGSound, UCF and ActivityNet) validate that the proposed AVMST outperforms existing state-of-the-art methods by a significant margin. The code and pre-trained models are available at https://github.com/liwr-hit/ICME23_AVMST.
Wenrui Li 0001, Zhengyu Ma, Liang-Jian Deng, Hengyu Man, Xiaopeng Fan 0001
ICME2
2023 Reservoir Computing Transformer for Image-Text Retrieval
abstract
Although the attention mechanism in transformers has proven successful in image-text retrieval tasks, most transformer models suffer from a large number of parameters. Inspired by brain circuits that process information with recurrent connected neurons, we propose a novel Reservoir Computing Transformer Reasoning Network (RCTRN) for image-text retrieval. The proposed RCTRN employs a two-step strategy to focus on feature representation and data distribution of different modalities respectively. Specifically, we send visual and textual features through a unified meshed reasoning module, which encodes multi-level feature relationships with prior knowledge and aggregates the complementary outputs in a more effective way. The reservoir reasoning network is proposed to optimize memory connections between features at different stages and address the data distribution mismatch problem introduced by the unified scheme. To investigate the significance of the low power dissipation and low bandwidth characteristics of RRN in practical scenarios, we deployed the model in the wireless transmission system, demonstrating that RRN's optimization of data structures also has a certain robustness against channel noise. Extensive experiments on two benchmark datasets, Flickr30K and MS-COCO, demonstrate the superiority of RCTRN in terms of performance and low-power dissipation compared to state-of-the-art baselines.
Wenrui Li 0001, Zhengyu Ma, Liang-Jian Deng, Penghong Wang, Jinqiao Shi, Xiaopeng Fan 0001
ACM Multimedia2
2023 Motion-Decoupled Spiking Transformer for Audio-Visual Zero-Shot Learning
abstract
Audio-visual zero-shot learning (ZSL) has attracted board attention, as it could classify video data from classes that are not observed during training. However, most of the existing methods are restricted to background scene bias and fewer motion details by employing a single-stream network to process scenes and motion information as a unified entity. In this paper, we address this challenge by proposing a novel dual-stream architecture Motion-Decoupled Spiking Transformer (MDFT) to explicitly decouple the contextual semantic information and highly sparsity dynamic motion information. Specifically, The Recurrent Joint Learning Unit (RJLU) could extract contextual semantic information effectively and understand the environment in which actions occur by capturing joint knowledge between different modalities. By converting RGB images to events, our approach effectively captures motion information while mitigating the influence of background scene biases, leading to more accurate classification results. We utilize the inherent strengths of Spiking Neural Networks (SNNs) to process highly sparsity event data efficiently. Additionally, we introduce a Discrepancy Analysis Block (DAB) to model the audio motion features. To enhance the efficiency of SNNs in extracting dynamic temporal and motion information, we dynamically adjust the threshold of Leaky Integrate-and-Fire (LIF) neurons based on the statistical cues of global motion and contextual semantic information. Our experiments demonstrate the effectiveness of MDFT, which consistently outperforms state-of-the-art methods across mainstream benchmarks. Moreover, we find that motion information serves as a powerful regularization for video networks, where using it improves the accuracy of HM and ZSL by 19.1% and 38.4%, respectively.
Wenrui Li 0001, Xi-Le Zhao, Zhengyu Ma, Xiaopeng Fan 0001, Yonghong Tian 0001
ACM Multimedia3
2023 Parallel Spiking Neurons with High Efficiency and Ability to Learn Long-term Dependencies
abstract
Vanilla spiking neurons in Spiking Neural Networks (SNNs) use charge-fire-reset neuronal dynamics, which can only be simulated serially and can hardly learn long-time dependencies. We find that when removing reset, the neuronal dynamics can be reformulated in a non-iterative form and parallelized. By rewriting neuronal dynamics without reset to a general formulation, we propose the Parallel Spiking Neuron (PSN), which generates hidden states that are independent of their predecessors, resulting in parallelizable neuronal dynamics and extremely high simulation speed. The weights of inputs in the PSN are fully connected, which maximizes the utilization of temporal information. To avoid the use of future inputs for step-by-step inference, the weights of the PSN can be masked, resulting in the masked PSN. By sharing weights across time-steps based on the masked PSN, the sliding PSN is proposed to handle sequences of varying lengths. We evaluate the PSN family on simulation speed and temporal/static data classification, and the results show the overwhelming advantage of the PSN family in efficiency and accuracy. To the best of our knowledge, this is the first study about parallelizing spiking neurons and can be a cornerstone for the spiking deep learning research. Our codes are available at https://github.com/fangwei123456/Parallel-Spiking-Neuron.
Wei Fang 0006, Zhaofei Yu, Zhaokun Zhou, Yanqi Chen, Zhengyu Ma, Timothée Masquelier, Yonghong Tian 0001
NeurIPS6
2023 The Style Transformer With Common Knowledge Optimization for Image-Text Retrieval
abstract
Image-text retrieval which associates different modalities has drawn broad attention due to its excellent research value and broad real-world application. However, most of the existing methods haven't taken the high-level semantic relationships (“style embedding”) and common knowledge from multi-modalities into full consideration. To this end, we introduce a novel style transformer network with common knowledge optimization (CKSTN) for image-text retrieval. The main module is the common knowledge adaptor (CKA) with both the style embedding extractor (SEE) and the common knowledge optimization (CKO) modules. Specifically, the SEE uses the sequential update strategy to effectively connect the features of different stages in SEE. The CKO module is introduced to dynamically capture the latent concepts of common knowledge from different modalities. Besides, to get generalized temporal common knowledge, we propose a sequential update strategy to effectively integrate the features of different layers in SEE with previous common feature units. CKSTN demonstrates the superiorities of the state-of-the-art methods in image-text retrieval on MSCOCO and Flickr30 K datasets. Moreover, CKSTN is constructed based on the lightweight transformer which is more convenient and practical for the application of real scenes, due to the better performance and lower parameters.
Wenrui Li 0001, Zhengyu Ma, Jinqiao Shi, Xiaopeng Fan 0001
IEEE Signal Process. Lett.2
2023 Neuron-Based Spiking Transmission and Reasoning Network for Robust Image-Text Retrieval
abstract
Most of the image-text retrieval methods carry out accurate results using fine-grained features for feature alignment. However, extracting the robustness features while maintaining the retrieval accuracy in wireless communication is still a challenge, especially with channel noises and limited transmission bandwidth. Inspired by spike signals of neurons in the human brain, we propose the neuron-based spiking transmission and reasoning network (NSTRN). In this way, the features are compressed into compacted efficient representations. In NSTRN, we construct the feature sender based on spiking activation function to selectively encode only important information in images and sentences into binary codes, and reduce the transmission cost. Moreover, the feature receiver is designed as a recurrent architecture and applies both temporal attention and global attention blocks to memorize long-term information. Finally, to compensate for the loss of visual concepts in transmission, we use the global textual features as coefficients to guide the formation of visual features in the training stage. The traditional CNN-based joint source-channel coding model outputs float-point encoded features, which requires additional quantization steps to convert features into binary bitstreams in the practical wireless communication system. Instead, the spiking neural networks (SNNs) directly use binary spike trains to reduce the computation complexity caused by the quantization steps. More importantly, SNNs can naturally encode the asynchronous event streams and inhibit the discrete noisy events to extract robust information. Even with binary bitstreams, NSTRN shows effectiveness compared with the state-of-the-art image-text retrieval methods. In the wireless communication scenario, NSTRN not only reduces the transmission bandwidth but also alleviates the “cliff effect” to a certain extent in the traditional separate encoding methods. To the best of our knowledge, this is the first work using SNNs on robust image-text retrieval.
Wenrui Li 0001, Zhengyu Ma, Liang-Jian Deng, Xiaopeng Fan 0001, Yonghong Tian 0001
IEEE Trans. Circuits Syst. Video Technol.2
2022 State Transition of Dendritic Spines Improves Learning of Sparse Spiking Neural Networks
abstract
Spiking Neural Networks (SNNs) are considered a promising alternative to Artificial Neural Networks (ANNs) for their event-driven computing paradigm when deployed on energy-efficient neuromorphic hardware. Recently, deep SNNs have shown breathtaking performance improvement through cutting-edge training strategy and flexible structure, which also scales up the number of parameters and computational burdens in a single network. Inspired by the state transition of dendritic spines in the filopodial model of spinogenesis, we model different states of SNN weights, facilitating weight optimization for pruning. Furthermore, the pruning speed can be regulated by using different functions describing the growing threshold of state transition. We organize these techniques as a dynamic pruning algorithm based on nonlinear reparameterization mapping from spine size to SNN weights. Our approach yields sparse deep networks on the large-scale dataset (SEW ResNet18 on ImageNet) while maintaining state-of-the-art low performance loss ( 3% at 88.8% sparsity) compared to existing pruning methods on directly trained SNNs. Moreover, we find out pruning speed regulation while learning is crucial to avoiding disastrous performance degradation at the final stages of training, which may shed light on future work on SNN pruning.
Yanqi Chen, Zhaofei Yu, Wei Fang 0006, Zhengyu Ma, Tiejun Huang 0001, Yonghong Tian 0001
ICML4