EDBT 2026 Demo / reviewers in the wild / expert
Qi Xu 0008
dblp:84/1680-8
· DBLP profile ↗
52ranked-venue papers
14as first author
46since 2021 · last 2026
0000-0001-9245-5544ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 37 · 12 first-author · 33 since 2021Graphics, computer vision, multimedia, augmented reality and games · 26 · 5 first-author · 25 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 4 since 2021Systems, architecture and hardware · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Spatial-Frequency Spiking Neural Network for Underwater Object DetectionabstractUnderwater object detection presents significant challenges due to the unique visual degradations in underwater environments, such as low contrast, poor visibility, and blurry object boundaries. While ANNs have achieved impressive detection accuracy, their high computational cost and power consumption limit their deployment in resource-constrained underwater platforms. In this work, we propose a Spatial-Frequency Spiking Neural Network (SFSNN) that combines the energy-efficient and event-driven nature of Spiking Neural Networks (SNNs) with the discriminative power of spatial-frequency analysis. SFSNN introduces a novel spatial-frequency spiking module that integrates spatial and frequency-domain representations, enhancing edge and texture features crucial for object detection in murky waters. Furthermore, we adapt the YOLOX architecture into a spike-based detector via ANN-to-SNN conversion using signed spiking neurons. Extensive experiments on the RUOD dataset demonstrate that SFSNN achieves superior performance over both SNN- and ANN-based detection models, offering a compelling solution for low-power underwater object detection. Long Chen 0019, Wei Miao 0006, Yunzhi Zhuge, Hongming Xu 0002, Qi Xu 0008 |
AAAI | 7 |
| 2026 | AVM: Towards Structure-Preserving Neural Response Modeling in the Visual Cortex Across Stimuli and IndividualsabstractWhile deep learning models have shown strong performance in simulating neural responses, they often fail to clearly separate stable visual encoding from condition-specific adaptation, which limits their ability to generalize across stimuli and individuals. We introduce the Adaptive Visual Model (AVM), a structure-preserving framework that enables condition-aware adaptation through modular subnetworks, without modifying the core representation. AVM keeps a Vision Transformer-based encoder frozen to capture consistent visual features, while independently trained modulation paths account for neural response variations driven by stimulus content and subject identity. We evaluate AVM in three experimental settings, including stimulus-level variation, cross-subject generalization, and cross-dataset adaptation, all of which involve structured changes in inputs and individuals. Across two large-scale mouse V1 datasets, AVM outperforms the state-of-the-art V1T model by approximately 2% in predictive correlation, demonstrating robust generalization, interpretable condition-wise modulation, and high architectural efficiency. Specifically, AVM achieves a 9.1% improvement in explained variance (FEVE) under the cross-dataset adaptation setting. These results suggest that AVM provides a unified framework for adaptive neural modeling across biological and experimental conditions, offering a scalable solution under structural constraints. Its design may inform future approaches to cortical modeling in both neuroscience and biologically inspired AI systems. Qi Xu 0008, Shuai Gong, Xuming Ran, Haihua Luo, Yangfan Hu |
AAAI | 1 |
| 2026 | Distillation-Guided Structural Transfer for Continual Learning Beyond Sparse Distributed MemoryabstractSparse neural systems are gaining traction for efficient continual learning due to their modularity and low interference. Architectures like Sparse Distributed Memory Multi-Layer Perceptrons (SDMLP) construct task-specific subnetworks via Top-K activation and have shown resilience against catastrophic forgetting. However, their rigid modularity poses two fundamental challenges: (1) the isolation of sparse subnetworks severely limits cross-task knowledge reuse; and (2) increased sparsity reduces interference but often degrades performance due to constrained feature sharing.We propose Selective Subnetwork Distillation (SSD), a structurally guided continual learning framework that treats distillation not as a regularizer, but as a topology-aligned information conduit. By identifying neurons with high activation frequency, SSD selectively distills knowledge within previous Top-K subnetworks and output logits—without requiring replay or task labels—preserving both sparsity and functional specialization.Unlike conventional distillation, SSD operates under hard modular constraints and enables structural realignment without altering the sparse architecture.While our method is validated on SDMLP, its structure-aligned mechanism has the potential to generalize to other sparse networks as a plug-in module for promoting representation sharing.Comprehensive experiments on Split CIFAR-10, CIFAR-100, and MNIST demonstrate that SSD improves accuracy, retention, and manifold coverage, offering a structurally grounded solution to sparse continual learning. Huiyan Xue, Xuming Ran, Qi Xu 0008, Enhui Li, Yi Xu 0008, Qiang Zhang 0008 |
AAAI | 4 |
| 2026 | Dual selective gleason pattern-aware multiple instance learning with uncertainty regularization for grade group prediction in histopathology images
Hongming Xu 0002, Qi Xu 0008, Ilkka Pölönen, Fengyu Cong |
Medical Image Anal. | 4 |
| 2026 | A complex-valued widening spiking neural network
Fang Liu 0017, Witold Pedrycz, Qi Xu 0008, Jialin Xu, Jie Yang 0007, Wei Wu 0010 |
Neural Networks | 3 |
| 2026 | Key-value pair-free continual learner via task-specific prompt-prototypeabstractContinual learning aims to enable models to acquire new knowledge while retaining previously learned information. Prompt-based methods have shown remarkable performance in this domain; however, they typically rely on key-value pairing, which can introduce inter-task interference and hinder scalability. To overcome these limitations, we propose a novel approach employing task-specific Prompt-Prototype (ProP), thereby eliminating the need for key-value pairs. In our method, task-specific prompts facilitate more effective feature learning for the current task, while corresponding prototypes capture the representative features of the input. During inference, predictions are generated by binding each task-specific prompt with its associated prototype. Additionally, we introduce regularization constraints during prompt initialization to penalize excessively large values, thereby enhancing stability. Experiments on several widely used datasets demonstrate the effectiveness of the proposed method. In contrast to mainstream prompt-based approaches, our framework removes the dependency on key-value pairs, offering a fresh perspective for future continual learning research. Haihua Luo, Xuming Ran, Zhengji Li, Huiyan Xue, Jiangrong Shen, Tommi Kärkkäinen, Qi Xu 0008, Fengyu Cong |
Neural Networks | 8 |
| 2026 | Context-Infused Trajectories: Enhancing Context and Frame Consistency in Reasoning Video Object SegmentationabstractReasoning video object segmentation (ReaVOS) aims to segment referred objects in video sequences based on implicit and complex linguistic queries. Existing methods typically compress limited video frames into pooled representations and prompt multimodal large language models (MLLMs) to generate a single global segmentation token. However, this strategy lacks explicit contextual guidance and causes substantial loss of spatial details, limiting capability and segmentation consistency. To overcome these limitations, we introduce Context-infused Consistent Video Segmentor (CiCVS), a novel framework leveraging contextual information to guide generation of temporally coherent and accurate mask trajectories. CiCVS incorporates a Hierarchical Frame Sampling (HFS) module, which globally samples support frames across the entire video to ensure broad temporal coverage, and then uniformly selects target frames within the support set. It also employs a Contextual Token Prompting (CTP) module, which utilizes contextual cues from support frames to guide the MLLM in generating specialized tokens for various target frames, enabling the model to capture intricate temporal patterns and ensure consistency across long-range sequences. At the core of CTP is the Multimodal Injection Compressor (MIC) block, which efficiently integrates support frame features and textual semantic information into a compact set of latent queries, enhancing temporal-level object perception. To further advance the ReaVOS field, we introduce the CoCoRVOS benchmark, which features more temporally intricate reasoning instructions and a diverse set of video scenarios. Extensive experiments demonstrate that CiCVS establishes a new state-of-the-art on multiple benchmarks, achieving significant improvements in $\mathcal {J}\& \mathcal {F}$ scores, including +2.7 on CoCoRVOS, +1.4 on ReVOS, and +7.0 on ReasonVOS, underscoring its superior contextual reasoning and segmentation capabilities. Yunzhi Zhuge, Sitong Gong, Lu Zhang 0053, Qi Xu 0008, Wenda Zhao 0003, Jin Zhan, Huchuan Lu |
IEEE Trans. Image Process. | 4 |
| 2026 | Generic-to-Personalised Learning for Multimodal Image Synthesis With Bidirectional Variational GANabstractMultimodal image synthesis, which predicts target-modality images from source-modality images, has garnered considerable attention in the field of clinical diagnosis. Both unidirectional and bidirectional multimodal image synthesis methods have been explored in the medical domain, however, unidirectional models heavily rely on paired images, while current bidirectional models typically overlook local image details due to their unsupervised training patterns. In this work, we propose a Bidirectional Variational Generative Adversarial Network (BVGAN) for multimodal image synthesis, which achieves high-quality bidirectional translations between any two modalities using only a limited number paired images. Firstly, BVGAN's generator incorporates a variational structure (VAS) to regularise the latent space for noise reduction. This regularisation imposes smoothness to the latent space, enabling BVGAN to produce high-quality, noise-free images. Secondly, a novel generic-to-personalised (GTP) learning strategy is introduced to train BVGAN and reduce its reliance on a large sets of paired images. GTP initially leverages an unsupervised learning model to capture the global mapping between two modalities using unpaired images from generic patients. It then applies a supervised learning model to refine the mapping for individual patient, enhancing image details. Finally, the GTP learning strategy along with VAS enables BVGAN to achieve state-of-the-art performance on two multi-modality medical datasets: Brain CTMRI and BRATS. Long Chen 0019, Xirui Dong, Jiangrong Shen, Lu Zhang 0053, Qi Xu 0008, Gang Pan 0001, Qiang Zhang 0008 |
IEEE Trans. Multim. | 5 |
| 2025 | Multi-View Incremental Learning with Structured Hebbian Plasticity for Enhanced Fusion EfficiencyabstractThe rapid evolution of multimedia technology has revolutionized human perception, paving the way for multi-view learning. However, traditional multi-view learning approaches are tailored for scenarios with fixed data views, falling short of emulating the intricate cognitive procedures of the human brain processing signals sequentially. Our cerebral architecture seamlessly integrates sequential data through intricate feed-forward and feedback mechanisms. In stark contrast, traditional methods struggle to generalize effectively when confronted with data spanning diverse domains, highlighting the need for innovative strategies that can mimic the brain's adaptability and dynamic integration capabilities. In this paper, we propose a bio-neurologically inspired multi-view incremental framework named MVIL aimed at emulating the brain's fine-grained fusion of sequentially arriving views. MVIL lies two fundamental modules: structured Hebbian plasticity and synaptic partition learning. The structured Hebbian plasticity reshapes the structure of weights to express the high correlation between view representations, facilitating a fine-grained fusion of view representations. Moreover, synaptic partition learning is efficient in alleviating drastic changes in weights and also retaining old knowledge by inhibiting partial synapses. These modules bionically play a central role in reinforcing crucial associations between newly acquired information and existing knowledge repositories, thereby enhancing the network's capacity for generalization. Experimental results on six benchmark datasets show MVIL's effectiveness over state-of-the-art methods. Ailin Song, Huifeng Yin, Shuai Zhong, Fuhai Chen, Qi Xu 0008, Shiping Wang, Mingkun Xu |
AAAI | 6 |
| 2025 | SpikingYOLOX: Improved YOLOX Object Detection with Fast Fourier Convolution and Spiking Neural NetworksabstractIn recent years, with the advancements in brain science, spiking neural networks (SNNs) have garnered significant attention. SNNs can generate spikes that mimic the function of neurons transmission in humans brain, thereby significantly reducing computational costs by the event-driven nature during training. While deep SNNs have shown impressive performance on classification tasks, they still face challenges in more complex tasks such as object detection. In this paper, we propose SpikingYOLOX, extending the structure of the original YOLOX by introducing signed spiking neurons and fast Fourier convolution (FFC). The designed ternary signed spiking neurons could generate three kinds of spikes to obtain more robust features in the deep layer of the backbone. Meanwhile, we integrate FFC with SNN modules to enhance object detection performance, because its global receptive field is beneficial to the object detection task. Extensive experiments demonstrate that the proposed SpikingYOLOX achieves state-of-the-art performance among other SNN-based object detection methods. Wei Miao 0006, Jiangrong Shen, Qi Xu 0008, Timo Hämäläinen 0002, Yi Xu 0008, Fengyu Cong |
AAAI | 3 |
| 2025 | ALADE-SNN: Adaptive Logit Alignment in Dynamically Expandable Spiking Neural Networks for Class Incremental LearningabstractInspired by the human brain's ability to adapt to new tasks without erasing prior knowledge, we develop spiking neural networks (SNNs) with dynamic structures for Class Incremental Learning (CIL). Our analytical experiments reveal that limited datasets introduce biases in logits distributions among tasks. Fixed features from frozen past-task extractors can cause overfitting and hinder the learning of new tasks. To address these challenges, we propose the ALADE-SNN framework, which includes adaptive logit alignment for balanced feature representation and OtoN suppression to manage weights mapping frozen old features to new classes during training, releasing them during fine-tuning. This approach dynamically adjusts the network architecture based on analytical observations, improving feature extraction and balancing performance between new and old tasks. Experiment results show that ALADE-SNN achieves an average incremental accuracy of 75.42 ± 0.74% on the CIFAR100-B0 dataset over 10 incremental steps. ALADE-SNN not only matches the performance of DNN-based methods but also surpasses state-of-the-art SNN-based continual learning algorithms. This advancement enhances continual learning in neuromorphic computing, offering a brain-inspired, energy-efficient solution for real-time data processing. Wenyao Ni, Jiangrong Shen, Qi Xu 0008, Huajin Tang |
AAAI | 3 |
| 2025 | BIG-FUSION: Brain-Inspired Global-Local Context Fusion Framework for Multimodal Emotion Recognition in ConversationsabstractConsidering the importance of capturing both global conversational topics and local speaker dependencies for multimodal emotion recognition in conversations, current approaches first utilize sequence models like Transformer to extract global context information, then apply Graph Neural Networks to model local speaker dependencies for local context information extraction, coupled with Graph Contrastive Learning (GCL) to enhance node representation learning. However, this sequential design introduces potential biases: the extracted global context information inevitably influences subsequent processing, compromising the independence and diversity of the original local features; current graph augmentation methods in GCL cannot consider both global and local context information in conversations to evaluate the node importance, hindering the learning of key information. Inspired by the human brain excels at handling complex tasks by efficiently integrating local and global information processing mechanisms, we propose an aligned global-local context fusion framework for sequence-based design to address these problems. This design includes a dual-attention Transformer and a dual-evaluation method for graph augmentation in GCL. The dual-attention Transformer combines global attention for overall context extraction with sliding-window attention for local context capture, both enhanced by spiking neuron dynamics. The dual-evaluation method in GCL comprises global importance evaluation to identify nodes crucial for overall conversation context, and local importance evaluation to detect nodes significant for local semantics, generating augmented graph views that preserve both global and local information. This approach ensures balanced information processing throughout the pipeline, enhancing biological plausibility and achieving superior emotion recognition. Yusong Wang 0003, Xuanye Fang, Huifeng Yin, Dongyuan Li, Qi Xu 0008, Yi Xu 0008, Shuai Zhong, Mingkun Xu |
AAAI | 6 |
| 2025 | FSTA-SNN: Frequency-Based Spatial-Temporal Attention Module for Spiking Neural NetworksabstractSpiking Neural Networks (SNNs) are emerging as a promising alternative to Artificial Neural Networks (ANNs) due to their inherent energy efficiency. Owing to the inherent sparsity in spike generation within SNNs, the in-depth analysis and optimization of intermediate output spikes are often neglected. This oversight significantly restricts the inherent energy efficiency of SNNs and diminishes their advantages in spatiotemporal feature extraction, resulting in a lack of accuracy and unnecessary energy expenditure. In this work, we analyze the inherent spiking characteristics of SNNs from both temporal and spatial perspectives. In terms of spatial analysis, we find that shallow layers tend to focus on learning vertical variations, while deeper layers gradually learn horizontal variations of features. Regarding temporal analysis, we observe that there is not a significant difference in feature learning across different time steps. This suggests that increasing the time steps has limited effect on feature learning. Based on the insights derived from these analyses, we propose a Frequency-based Spatial-Temporal Attention (FSTA) module to enhance feature learning in SNNs. This module aims to improve the feature learning capabilities by suppressing redundant spike features. The experimental results indicate that the introduction of the FSTA module significantly reduces the spike firing rate of SNNs, demonstrating superior performance compared to state-of-the-art baselines across multiple datasets. Kairong Yu, Tianqing Zhang, Hongwei Wang 0001, Qi Xu 0008 |
AAAI | 4 |
| 2025 | ODA-GAN: Orthogonal Decoupling Alignment GAN Assisted by Weakly-supervised Learning for Virtual Immunohistochemistry StainingabstractRecently, virtual staining has emerged as a promising alternative to revolutionize histological staining by digitally generating stains. However, most existing methods suffer from the curse of staining unreality and unreliability. In this paper, we propose the Orthogonal Decoupling Alignment Generative Adversarial Network (ODA-GAN) for unpaired virtual immunohistochemistry (IHC) staining. Our approach is based on the assumption that an image consists of IHC staining-related features, which influence staining distribution and intensity, and staining-unrelated features, such as tissue morphology. Leveraging a pathology foundation model, we first develop a weakly-supervised segmentation pipeline as an alternative to expert annotations. We introduce an Orthogonal MLP (O-MLP) module to project image features into an orthogonal space, decoupling them into staining-related and unrelated components. Additionally, we propose a Dual-stream PatchNCE (DPNCE) loss to resolve contrastive learning contradictions in the staining-related space, thereby enhancing staining accuracy. To further improve realism, we introduce a Multi-layer Domain Alignment (MDA) module to bridge the domain gap between generated and real IHC images. Evaluations on three benchmark datasets show that our ODA-GAN reaches state-of-the-art (SOTA) performance. Our source code is available at https://github.com/ittong/ODA-GAN. Mingkang Wang, Zhongze Wang, Hongkai Wang 0002, Qi Xu 0008, Fengyu Cong, Hongming Xu 0002 |
CVPR | 5 |
| 2025 | Temporal Separation with Entropy Regularization for Knowledge Distillation in Spiking Neural NetworksabstractSpiking Neural Networks (SNNs), inspired by the human brain, offer significant computational efficiency through discrete spike-based information transfer. Despite their potential to reduce inference energy consumption, a performance gap persists between SNNs and Artificial Neural Networks (ANNs), primarily due to current training methods and inherent model limitations. While recent research has aimed to enhance SNN learning by employing knowledge distillation (KD) from ANN teacher networks, traditional distillation techniques often overlook the distinctive spatiotemporal properties of SNNs, thus failing to fully leverage their advantages. To overcome these challenge, we propose a novel logit distillation method characterized by temporal separation and entropy regularization. This approach improves existing SNN distillation techniques by performing distillation learning on logits across different time steps, rather than merely on aggregated output features. Furthermore, the integration of entropy regularization stabilizes model optimization and further boosts the performance. Extensive experimental results indicate that our method surpasses prior SNN distillation strategies, whether based on logit distillation, feature distillation, or a combination of both. Our project is available at https://github.com/yukairong/TSER. Kairong Yu, Chengting Yu, Tianqing Zhang, Xiaochen Zhao, Hongwei Wang 0001, Qiang Zhang 0008, Qi Xu 0008 |
CVPR | 8 |
| 2025 | STAA-SNN: Spatial-Temporal Attention Aggregator for Spiking Neural NetworksabstractSpiking Neural Networks (SNNs) have gained significant attention due to their biological plausibility and energy efficiency, making them promising alternatives to Artificial Neural Networks (ANNs). However, the performance gap between SNNs and ANNs remains a substantial challenge hindering the widespread adoption of SNNs. In this paper, we propose a Spatial-Temporal Attention Aggregator SNN (STAA-SNN) framework, which dynamically focuses on and captures both spatial and temporal dependencies. First, we introduce a spike-driven self-attention mechanism specifically designed for SNNs. Additionally, we pioneeringly incorporate position encoding to integrate latent temporal relationships into the incoming features. For spatial-temporal information aggregation, we employ step attention to selectively amplify relevant features to variant steps. Finally, we implement a time-step random dropout strategy to avoid local optima. The framework demonstrates exceptional performance across diverse datasets and exhibits strong generalization capabilities. Notably, STAA-SNN achieves state-of-the-art results on neuromorphic datasets CIFAR10-DVS of 82.10% and with performances of 97.14%, 82.05% and 70.40% on the static datasets CIFAR-10, CIFAR-100 and ImageNet, respectively. Furthermore, this model exhibits improved performance ranging from 0.33% to 2.80% with fewer time steps. Tianqing Zhang, Kairong Yu, Xian Zhong, Hongwei Wang 0001, Qi Xu 0008, Qiang Zhang 0008 |
CVPR | 5 |
| 2025 | Improving the Sparse Structure Learning of Spiking Neural Networks from the View of Compression EfficiencyabstractThe human brain utilizes spikes for information transmission and dynamically reorganizes its network structure to boost energy efficiency and cognitive capabilities throughout its lifespan. Drawing inspiration from this spike-based computation, Spiking Neural Networks (SNNs) have been developed to construct event-driven models that emulate this efficiency. Despite these advances, deep SNNs continue to suffer from over-parameterization during training and inference, a stark contrast to the brain’s ability to self-organize. Furthermore, existing sparse SNNs are challenged by maintaining optimal pruning levels due to a static pruning ratio, resulting in either under or over-pruning.
In this paper, we propose a novel two-stage dynamic structure learning approach for deep SNNs, aimed at maintaining effective sparse training from scratch while optimizing compression efficiency.
The first stage evaluates the compressibility of existing sparse subnetworks within SNNs using the PQ index, which facilitates an adaptive determination of the rewiring ratio for synaptic connections based on data compression insights. In the second stage, this rewiring ratio critically informs the dynamic synaptic connection rewiring process, including both pruning and regrowth. This approach significantly improves the exploration of sparse structures training in deep SNNs, adapting sparsity dynamically from the point view of compression efficiency.
Our experiments demonstrate that this sparse training approach not only aligns with the performance of current deep SNNs models but also significantly improves the efficiency of compressing sparse SNNs. Crucially, it preserves the advantages of initiating training with sparse models and offers a promising solution for implementing Edge AI on neuromorphic hardware. Jiangrong Shen, Qi Xu 0008, Gang Pan 0001, Badong Chen |
ICLR | 2 |
| 2025 | Hybrid Spiking Vision Transformer for Object Detection with Event CamerasabstractEvent-based object detection has attracted increasing attention for its high temporal resolution, wide dynamic range, and asynchronous address-event representation. Leveraging these advantages, spiking neural networks (SNNs) have emerged as a promising approach, offering low energy consumption and rich spatiotemporal dynamics. To further enhance the performance of event-based object detection, this study proposes a novel hybrid spike vision Transformer (HsVT) model. The HsVT model integrates a spatial feature extraction module to capture local and global features, and a temporal feature extraction module to model time dependencies and long-term patterns in event sequences. This combination enables HsVT to capture spatiotemporal features, improving its capability in handling complex event-based object detection tasks. To support research in this area, we developed the Fall Detection dataset as a benchmark for event-based object detection tasks. The Fall DVS detection dataset protects facial privacy and reduces memory usage thanks to its event-based representation. Experimental results demonstrate that HsVT outperforms existing SNN methods and achieves competitive performance compared to ANN-based models, with fewer parameters and lower energy consumption. Qi Xu 0008, Jiangrong Shen, Biwu Chen, Huajin Tang, Gang Pan 0001 |
ICML | 1 |
| 2025 | Self-cross Feature based Spiking Neural Networks for Efficient Few-shot LearningabstractDeep neural networks (DNNs) excel in computer vision tasks, especially, few-shot learning (FSL), which is increasingly important for generalizing from limited examples. However, DNNs are computationally expensive with scalability issues in real world. Spiking Neural Networks (SNNs), with their event-driven nature and low energy consumption, are particularly efficient in processing sparse and dynamic data, though they still encounter difficulties in capturing complex spatiotemporal features and performing accurate cross-class comparisons. To further enhance the performance and efficiency of SNNs in few-shot learning, we propose a few-shot learning framework based on SNNs, which combines a self-feature extractor module and a cross-feature contrastive module to refine feature representation and reduce power consumption. We apply the combination of temporal efficient training loss and InfoNCE loss to optimize the temporal dynamics of spike trains and enhance the discriminative power. Experimental results show that the proposed FSL-SNN significantly improves the classification performance on the neuromorphic dataset N-Omniglot, and also achieves competitive performance to ANNs on static datasets such as CUB and miniImageNet with low power consumption. Qi Xu 0008, Junyang Zhu, Dongdong Zhou, Jiangrong Shen, Qiang Zhang 0008 |
ICML | 1 |
| 2025 | Efficient ANN-SNN Conversion with Error Compensation LearningabstractArtificial neural networks (ANNs) have demonstrated outstanding performance in numerous tasks, but deployment in resource-constrained environments remains a challenge due to their high computational and memory requirements. Spiking neural networks (SNNs) operate through discrete spike events and offer superior energy efficiency, providing a bio-inspired alternative. However, current ANN-to-SNN conversion often results in significant accuracy loss and increased inference time due to conversion errors such as clipping, quantization, and uneven activation. This paper proposes a novel ANN-to-SNN conversion framework based on error compensation learning. We introduce a learnable threshold clipping function, dual-threshold neurons, and an optimized membrane potential initialization strategy to mitigate the conversion error. Together, these techniques address the clipping error through adaptive thresholds, dynamically reduce the quantization error through dual-threshold neurons, and minimize the non-uniformity error by effectively managing the membrane potential. Experimental results on CIFAR-10, CIFAR-100, ImageNet datasets show that our method achieves high-precision and ultra-low latency among existing conversion methods. Using only two time steps, our method significantly reduces the inference time while maintains competitive accuracy of 94.75% on CIFAR-10 dataset under ResNet-18 structure. This research promotes the practical application of SNNs on low-power hardware, making efficient real-time processing possible. Chang Liu 0030, Jiangrong Shen, Xuming Ran, Mingkun Xu, Qi Xu 0008, Yi Xu 0008, Gang Pan 0001 |
ICML | 5 |
| 2025 | Enhancing Graph Contrastive Learning for Protein Graphs from Perspective of InvarianceabstractGraph Contrastive Learning (GCL) improves Graph Neural Network (GNN)-based protein representation learning by enhancing its generalization and robustness. Existing GCL approaches for protein representation learning rely on 2D topology, where graph augmentation is solely based on topological features, ignoring the intrinsic biological properties of proteins. Besides, 3D structure-based protein graph augmentation remains unexplored, despite proteins inherently exhibiting 3D structures. To bridge this gap, we propose novel biology-aware graph augmentation strategies from the perspective of invariance and integrate them into the protein GCL framework. Specifically, we introduce Functional Community Invariance (FCI)-based graph augmentation, which employs spectral constraints to preserve topology-driven community structures while incorporating residue-level chemical similarity as edge weights to guide edge sampling and maintain functional communities. Furthermore, we propose 3D Protein Structure Invariance (3-PSI)-based graph augmentation, leveraging dihedral angle perturbations and secondary structure rotations to retain critical 3D structural information of proteins while diversifying graph views. Extensive experiments on four different protein-related tasks demonstrate the superiority of our proposed GCL protein representation learning framework. Yusong Wang 0003, Shiyin Tan, Jialun Shen, Haobo Song, Qi Xu 0008, Prayag Tiwari, Mingkun Xu |
ICML | 6 |
| 2025 | TS-SNN: Temporal Shift Module for Spiking Neural NetworksabstractSpiking Neural Networks (SNNs) are increasingly recognized for their biological plausibility and energy efficiency, positioning them as strong alternatives to Artificial Neural Networks (ANNs) in neuromorphic computing applications. SNNs inherently process temporal information by leveraging the precise timing of spikes, but balancing temporal feature utilization with low energy consumption remains a challenge. In this work, we introduce Temporal Shift module for Spiking Neural Networks (TS-SNN), which incorporates a novel Temporal Shift (TS) module to integrate past, present, and future spike features within a single timestep via a simple yet effective shift operation. A residual combination method prevents information loss by integrating shifted and original features. The TS module is lightweight, requiring only one additional learnable parameter, and can be seamlessly integrated into existing architectures with minimal additional computational cost. TS-SNN achieves state-of-the-art performance on benchmarks like CIFAR-10 (96.72%), CIFAR-100 (80.28%), and ImageNet (70.61%) with fewer timesteps, while maintaining low energy consumption. This work marks a significant step forward in developing efficient and accurate SNN architectures. Kairong Yu, Tianqing Zhang, Qi Xu 0008, Gang Pan 0001, Hongwei Wang 0001 |
ICML | 3 |
| 2025 | Dual Selective Gleason Pattern-Aware Multiple Instance Learning for Grade Group Prediction in Histopathology Images
Hongming Xu 0002, Qibin Zhang, Qi Xu 0008, Ilkka Pölönen, Fengyu Cong |
MICCAI (15) | 4 |
| 2025 | Predicting Radiation Therapy Response Based on Dynamic Temporal Feature Difference Fusion from Longitudinal MRI
Hongming Xu 0002, Qibin Zhang, Qi Xu 0008, Ilkka Pölönen, Fengyu Cong |
MICCAI (16) | 4 |
| 2025 | Advanced SpikingYOLOX: Extending Spiking Neural Network on Object Detection with Spike-based Partial Self-Attention and 2D-Spiking TransformerabstractBrain-inspired Spiking Neural Networks (SNNs) have garnered significant attention due to their bio-plausibility and low power consumption advantages compared to Artificial Neural Networks (ANNs). However, the application of SNN in computer vision remains limited, primarily due to their inferior performance. In this work, we aim to bridge the performance gap between ANNs and SNNs in object detection by our Advanced SpikingYOLOX. The proposed approach extends the SpikingYOLOX with two key innovations: PSA-SNN and 2D-Spiking Transformer, both designed to enhance object detection performance. PSA-SNN extends spike-based self-attention by incorporating high-speed partial self-attention with an SNN-based 2D-Spiking Transformer in the deepest layer of the backbone, significantly improving feature extraction. The 2D-Spiking Transformer redefines the role of spiking neurons in Transformer sequences (Key, Query, Value), demonstrating that applying an additional spiking layer solely to the Value sequence yields the best performance while maintaining computational efficiency in spike-driven Transformers. We conduct extensive experiments on static images and the Advanced SpikingYOLOX achieves state-of-the-art performance among other SNN-based object detection methods. This work paves the way for more advanced SNN applications in object detection and broader computer vision tasks. Wei Miao 0006, Jiangrong Shen, Hongming Xu 0002, Tommi Kärkkäinen, Qi Xu 0008, Yi Xu 0008, Fengyu Cong |
ACM Multimedia | 5 |
| 2025 | Spiking Neural Networks with Temporal Attention-Guided Adaptive Fusion for imbalanced Multi-modal Learning
Jiangrong Shen, Yulin Xie, Qi Xu 0008, Gang Pan 0001, Huajin Tang, Badong Chen |
ACM Multimedia | 3 |
| 2025 | Context gating in spiking neural networks: Achieving lifelong learning through integration of local and global plasticity
Jiangrong Shen, Wenyao Ni, Qi Xu 0008, Gang Pan 0001, Huajin Tang |
Knowl. Based Syst. | 3 |
| 2025 | Temporal spiking generative adversarial networks for heading direction decoding
Jiangrong Shen, Jian K. Liu, Qi Xu 0008, Gang Pan 0001, Xiaodong Chen 0005, Huajin Tang |
Neural Networks | 5 |
| 2025 | Multi-Task Adaptive Resolution Network for Lymph Node Metastasis Diagnosis From Whole Slide Images of Colorectal CancerabstractAutomated detection of lymph node metastasis (LNM) holds great potential to alleviate the workload of doctors and reduce misinterpretations. Despite the practical successes achieved, effectively addressing the highly complex and heterogeneous tumor microenvironment remains an open and challenging problem, especially when tumor subtypes intermingle and are difficult to delineate. In this paper, we propose a multi-task adaptive resolution network, named MAR-Net, for LNM detection and subtyping in complex mixed-type cancers. Specifically, we construct a resolution-aware module to mine heterogeneous diagnostic information, which exploits the multi-scale pyramid information and adaptively combines multi-resolution structured features for comprehensive representation. Additionally, we adopt a multi-task learning approach that simultaneously addresses LNM detection and subtyping, reducing model instability during optimization and improving performance across both tasks. More importantly, to rectify the potential misclassification of tumor subtypes, we elaborately design a hierarchical subtying refinement (HSR) algorithm that leverages a generic segmentation model informed by pathologists' prior knowledge. Evaluations have been conducted on three private and one public cancer datasets (554 WSIs, 4.8 million patches). Our experimental results demonstrate that the proposed method consistently achieves superior performance compared to the state-of-the-art methods, achieving 0.5% to 3.2% higher AUC in LNM detection and 3.8% to 4.4% higher AUC in LNM subtyping. Su-Jin Shin, Mingkang Wang, Qi Xu 0008, Guiyang Jiang, Fengyu Cong, Jeonghyun Kang, Hongming Xu 0002 |
IEEE J. Biomed. Health Informatics | 4 |
| 2025 | Robust Sensory Information Reconstruction and Classification With Augmented SpikesabstractSensory information recognition is primarily processed through the ventral and dorsal visual pathways in the primate brain visual system, which exhibits layered feature representations bearing a strong resemblance to convolutional neural networks (CNNs), encompassing reconstruction and classification. However, existing studies often treat these pathways as distinct entities, focusing individually on pattern reconstruction or classification tasks, overlooking a key feature of biological neurons, the fundamental units for neural computation of visual sensory information. Addressing these limitations, we introduce a unified framework for sensory information recognition with augmented spikes. By integrating pattern reconstruction and classification within a single framework, our approach not only accurately reconstructs multimodal sensory information but also provides precise classification through definitive labeling. Experimental evaluations conducted on various datasets including video scenes, static images, dynamic auditory scenes, and functional magnetic resonance imaging (fMRI) brain activities demonstrate that our framework delivers state-of-the-art pattern reconstruction quality and classification accuracy. The proposed framework enhances the biological realism of multimodal pattern recognition models, offering insights into how the primate brain visual system effectively accomplishes the reconstruction and classification tasks through the integration of ventral and dorsal pathways. Qi Xu 0008, Sibo Liu, Xuming Ran, Jiangrong Shen, Huajin Tang, Jian K. Liu, Gang Pan 0001, Qiang Zhang 0008 |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2024 | Efficient Spiking Neural Networks with Sparse Selective Activation for Continual LearningabstractThe next generation of machine intelligence requires the capability of continual learning to acquire new knowledge without forgetting the old one while conserving limited computing resources. Spiking neural networks (SNNs), compared to artificial neural networks (ANNs), have more characteristics that align with biological neurons, which may be helpful as a potential gating function for knowledge maintenance in neural networks. Inspired by the selective sparse activation principle of context gating in biological systems, we present a novel SNN model with selective activation to achieve continual learning. The trace-based K-Winner-Take-All (K-WTA) and variable threshold components are designed to form the sparsity in selective activation in spatial and temporal dimensions of spiking neurons, which promotes the subpopulation of neuron activation to perform specific tasks. As a result, continual learning can be maintained by routing different tasks via different populations of neurons in the network. The experiments are conducted on MNIST and CIFAR10 datasets under the class incremental setting. The results show that the proposed SNN model achieves competitive performance similar to and even surpasses the other regularization-based methods deployed under traditional ANNs. Jiangrong Shen, Wenyao Ni, Qi Xu 0008, Huajin Tang |
AAAI | 3 |
| 2024 | Adaptive deep spiking neural network with global-local learning via balanced excitatory and inhibitory mechanismabstractThe training method of Spiking Neural Networks (SNNs) is an essential problem, and how to integrate local and global learning is a worthy research interest. However, the current integration methods do not consider the network conditions suitable for local and global learning, and thus fail to balance their advantages. In this paper, we propose an Excitation-Inhibition Mechanism-assisted Hybrid Learning(EIHL) algorithm that adjusts the network connectivity by using the excitation-inhibition mechanism and then switches between local and global learning according to the network connectivity. The experimental results on CIFAR10/100 and DVS-CIFAR10 demonstrate that the EIHL not only has better accuracy performance than other methods but also has excellent sparsity advantage. Especially, the Spiking VGG11 is trained by EIHL, STBP, and STDP on DVS_CIFAR10, respectively. The accuracy of the Spiking VGG11 model on EIHL is 62.45%, which is 4.35% higher than STBP and 11.40% higher than STDP, and the sparsity is 18.74%, which is 18.74% higher than the other two methods. Moreover, the excitation-inhibition mechanism used in our method also offers a new perspective on the field of SNN learning. Qi Xu 0008, Xuming Ran, Jiangrong Shen, Pan Lv, Qiang Zhang 0008, Gang Pan 0001 |
ICLR | 2 |
| 2024 | The Balanced Multi-Modal Spiking Neural Networks with Online Loss Adjustment and Time AlignmentabstractOptimizing multi-modal learning of SNNs has the advantages of energy efficiency and performance improvements. However, modality imbalance in multi-modal SNNs results in performance decline due to heterogeneity of modalities and temporal inconsistencies across different SNNs branches. In this paper, we propose the Balanced Multi-modal SNNs (BM-SNNs) model, equipped with a novel online loss adjustment (LA) algorithm and time alignment (TA) modules, ultimately achieving balanced training across multiple modalities. LA supervises the learning of uni-modal feature extractors by adding unimodal loss components without additional classifier. Moreover, the modulation factors enable the adaptive adjustment of unimodal learning rate. Furthermore, the proposed TA adopts the optimal timestep for different modalities to avoid information redundancy. Experimental results reveal that BM-SNNs model improves the performance of both multi-modal model and unimodal branches through adaptively exploiting intra-modal and cross-modal information. Jianing Han, Jiangrong Shen, Qi Xu 0008, Jian K. Liu, Huajin Tang |
ICME | 3 |
| 2024 | Towards efficient deep spiking neural networks construction with spiking activity based pruningabstractThe emergence of deep and large-scale spiking neural networks (SNNs) exhibiting high performance across diverse complex datasets has led to a need for compressing network models due to the presence of a significant number of redundant structural units, aiming to more effectively leverage their low-power consumption and biological interpretability advantages. Currently, most model compression techniques for SNNs are based on unstructured pruning of individual connections, which requires specific hardware support. Hence, we propose a structured pruning approach based on the activity levels of convolutional kernels named Spiking Channel Activity-based (SCA) network pruning framework. Inspired by synaptic plasticity mechanisms, our method dynamically adjusts the network’s structure by pruning and regenerating convolutional kernels during training, enhancing the model’s adaptation to the current target task. While maintaining model performance, this approach refines the network architecture, ultimately reducing computational load and accelerating the inference process. This indicates that structured dynamic sparse learning methods can better facilitate the application of deep SNNs in low-power and high-efficiency scenarios. Qi Xu 0008, Jiangrong Shen, Hongming Xu 0002, Long Chen 0019, Gang Pan 0001 |
ICML | 2 |
| 2024 | RSNN: Recurrent Spiking Neural Networks for Dynamic Spatial-Temporal Information Processing
Qi Xu 0008, Xuanye Fang, Jiangrong Shen, De Ma, Yi Xu 0008, Gang Pan 0001 |
ACM Multimedia | 1 |
| 2024 | Reversing Structural Pattern Learning with Biologically Inspired Knowledge Distillation for Spiking Neural NetworksabstractSpiking neural networks (SNNs) have superb characteristics in sensory information recognition tasks due to their biological plausibility. However, the performance of some current spiking-based models is limited by their structures which means either fully connected or too-deep structures bring too much redundancy. This redundancy from both connection and neurons is one of the key factors hindering the practical application of SNNs. Although Some pruning methods were proposed to tackle this problem, they normally ignored the fact the neural topology in the human brain could be adjusted dynamically. Inspired by this, this paper proposed an evolutionary-based structure construction method for constructing more reasonable SNNs. By integrating the knowledge distillation and connection pruning method, the synaptic connections in SNNs can be optimized dynamically to reach an optimal state. As a result, the structure of SNNs could not only absorb knowledge from the teacher model but also search for deep but sparse network topology. Experimental results on CIFAR100, Tiny-imagenet and DVS-Gesture show that the proposed structure learning method can get pretty well performance while reducing the connection redundancy. The proposed method explores a novel dynamical way for structure learning from scratch in SNNs which could build a bridge to close the gap between deep learning and bio-inspired neural dynamics. Qi Xu 0008, Xuanye Fang, Jiangrong Shen, Qiang Zhang 0008, Gang Pan 0001 |
ACM Multimedia | 1 |
| 2024 | CWSCNet: Channel-Weighted Skip Connection Network for Underwater Object DetectionabstractAutonomous underwater vehicles (AUVs) equipped with the intelligent underwater object detection technique is of great significance for underwater navigation. Advanced underwater object detection frameworks adopt skip connections to enhance the feature representation which further boosts the detection precision. However, we reveal two limitations of standard skip connections: 1) standard skip connections do not consider the feature heterogeneity, resulting in a sub-optimal feature fusion strategy; 2) feature redundancy exists in the skip connected features that not all the channels in the fused feature maps are equally important, the network learning should focus on the informative channels rather than the redundant ones. In this paper, we propose a novel channel-weighted skip connection network (CWSCNet) to learn multiple hyper fusion features for improving multi-scale underwater object detection. In CWSCNet, a novel feature fusion module, named channel-weighted skip connection (CWSC), is proposed to adaptively adjust the importance of different channels during feature fusion. The CWSC module removes feature heterogeneity that strengthens the compatibility of different feature maps, it also works as an effective feature selection strategy that enables CWSCNet to focus on learning channels with more object-related information. Extensive experiments on three underwater object detection datasets RUOD, URPC2017 and URPC2018 show that the proposed CWSCNet achieves comparable or state-of-the-art performances in underwater object detection. Long Chen 0019, Yunzhou Xie, Qi Xu 0008, Junyu Dong |
IEEE Trans. Image Process. | 4 |
| 2024 | Event-Assisted Blurriness Representation Learning for Blurry Image UnfoldingabstractThe goal of blurry image deblurring and unfolding task is to recover a single sharp frame or a sequence from a blurry one. Recently, its performance is greatly improved with introduction of a bio-inspired visual sensor, event camera. Most existing event-assisted deblurring methods focus on the design of powerful network architectures and effective training strategy, while ignoring the role of blur modeling in removing various blur in dynamic scenes. In this work, we propose to implicitly model blur in an image by computing blurriness representation with an event-assisted blurriness encoder. The learning of blurriness representation is formulated as a ranking problem based on specially synthesized pairs. Blurriness-aware image unfolding is achieved by integrating blur relevant information contained in the representation into a base unfolding network. The integration is mainly realized by the proposed blurriness-guided modulation and multi-scale aggregation modules. Experiments on GOPRO and HQF datasets show favorable performance of the proposed method against state-of-the-art approaches. More results on real-world data validate its effectiveness in recovering a sequence of latent sharp frames from a blurry image. Hao Ju 0004, Lei Yu 0006, Weihua He, Yaoyuan Wang, Qi Xu 0008, Shengming Li, Dong Wang 0004, Huchuan Lu, Xu Jia 0012 |
IEEE Trans. Image Process. | 7 |
| 2024 | Hierarchical Spiking-Based Model for Efficient Image Classification With Enhanced Feature Extraction and EncodingabstractThanks to their event-driven nature, spiking neural networks (SNNs) are surmised to be great computation-efficient models. The spiking neurons encode beneficial temporal facts and possess excessive anti-noise properties. However, the high-quality encoding of spatio-temporal complexity and also its training optimization of SNNs are restricted by means of the contemporary problem, this article proposes a novel hierarchical event-driven visual device to explore how information transmits and signifies in the retina the usage of biologically manageable mechanisms. This cognitive model is an augmented spiking-based framework consisting of the function learning capacity of convolutional neural networks (CNNs) with the cognition capability of SNNs. Furthermore, this visual device is modeled in a biological realism way with unsupervised learning rules and advanced spike firing rate encoding methods. We train and test them on some image datasets (Modified National Institute of Standards and Technology (MNIST), Canadian Institute for Advanced Research (CIFAR)10, and its noisy versions) to show that our mannequin can process greater vital data than present cognitive models. This article also proposes a novel quantization approach to make the proposed spiking-based model more efficient for neuromorphic hardware implementation. The outcomes show this joint CNN-SNN model can reap excessive focus accuracy and get more effective generalization ability. Qi Xu 0008, Jiangrong Shen, Jian K. Liu, Huajin Tang, Gang Pan 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2023 | ESL-SNNs: An Evolutionary Structure Learning Strategy for Spiking Neural NetworksabstractSpiking neural networks (SNNs) have manifested remarkable advantages in power consumption and event-driven property during the inference process. To take full advantage of low power consumption and improve the efficiency of these models further, the pruning methods have been explored to find sparse SNNs without redundancy connections after training. However, parameter redundancy still hinders the efficiency of SNNs during training. In the human brain, the rewiring process of neural networks is highly dynamic, while synaptic connections maintain relatively sparse during brain development. Inspired by this, here we propose an efficient evolutionary structure learning (ESL) framework for SNNs, named ESL-SNNs, to implement the sparse SNN training from scratch. The pruning and regeneration of synaptic connections in SNNs evolve dynamically during learning, yet keep the structural sparsity at a certain level. As a result, the ESL-SNNs can search for optimal sparse connectivity by exploring all possible parameters across time. Our experiments show that the proposed ESL-SNNs framework is able to learn SNNs with sparse structures effectively while reducing the limited accuracy. The ESL-SNNs achieve merely 0.28% accuracy loss with 10% connection density on the DVS-Cifar10 dataset. Our work presents a brand-new approach for sparse training of SNNs from scratch with biologically plausible evolutionary mechanisms, closing the gap in the expressibility between sparse training and dense training. Hence, it has great potential for SNN lightweight training and inference with low power consumption and small memory usage. Jiangrong Shen, Qi Xu 0008, Jian K. Liu, Yueming Wang 0001, Gang Pan 0001, Huajin Tang |
AAAI | 2 |
| 2023 | Constructing Deep Spiking Neural Networks from Artificial Neural Networks with Knowledge DistillationabstractSpiking neural networks (SNNs) are well-known as brain-inspired models with high computing efficiency, due to a key component that they utilize spikes as information units, close to the biological neural systems. Although spiking based models are energy efficient by taking advantage of discrete spike signals, their performance is limited by current network structures and their training methods. As discrete signals, typical SNNs cannot apply the gradient descent rules directly into parameter adjustment as artificial neural networks (ANNs). Aiming at this limitation, here we propose a novel method of constructing deep SNN models with knowledge distillation (KD) that uses ANN as the teacher model and SNN as the student model. Through the ANN-SNN joint training algorithm, the student SNN model can learn rich feature information from the teacher ANN model through the KD method, yet it avoids training SNN from scratch when communicating with non-differentiable spikes. Our method can not only build a more efficient deep spiking structure feasibly and reasonably but use few time steps to train the whole model compared to direct training or ANN to SNN methods. More importantly, it has a superb ability of noise immunity for various types of artificial noises and natural signals. The proposed novel method provides efficient ways to improve the performance of SNN through constructing deeper structures in a high-throughput fashion, with potential usage for light and efficient brain-inspired computing of practical scenarios. Qi Xu 0008, Jiangrong Shen, Jian K. Liu, Huajin Tang, Gang Pan 0001 |
CVPR | 1 |
| 2023 | EICIL: Joint Excitatory Inhibitory Cycle Iteration Learning for Deep Spiking Neural NetworksabstractSpiking neural networks (SNNs) have undergone continuous development and extensive study for decades, leading to increased biological plausibility and optimal energy efficiency. However, traditional training methods for deep SNNs have some limitations, as they rely on strategies such as pre-training and fine-tuning, indirect coding and reconstruction, and approximate gradients. These strategies lack a complete training model and require gradient approximation. To overcome these limitations, we propose a novel learning method named Joint Excitatory Inhibitory Cycle Iteration learning for Deep Spiking Neural Networks (EICIL) that integrates both excitatory and inhibitory behaviors inspired by the signal transmission of biological neurons.By organically embedding these two behavior patterns into one framework, the proposed EICIL significantly improves the bio-mimicry and adaptability of spiking neuron models, as well as expands the representation space of spiking neurons. Extensive experiments based on EICIL and traditional learning methods demonstrate that EICIL outperforms traditional methods on various datasets, such as CIFAR10 and CIFAR100, revealing the crucial role of the learning approach that integrates both behaviors during training. Zihang Shao, Xuanye Fang, Chaoran Feng 0001, Jiangrong Shen, Qi Xu 0008 |
NeurIPS | 6 |
| 2023 | Enhancing Adaptive History Reserving by Spiking Convolutional Block Attention Module in Recurrent Neural NetworksabstractSpiking neural networks (SNNs) serve as one type of efficient model to process spatio-temporal patterns in time series, such as the Address-Event Representation data collected from Dynamic Vision Sensor (DVS). Although convolutional SNNs have achieved remarkable performance on these AER datasets, benefiting from the predominant spatial feature extraction ability of convolutional structure, they ignore temporal features related to sequential time points. In this paper, we develop a recurrent spiking neural network (RSNN) model embedded with an advanced spiking convolutional block attention module (SCBAM) component to combine both spatial and temporal features of spatio-temporal patterns. It invokes the history information in spatial and temporal channels adaptively through SCBAM, which brings the advantages of efficient memory calling and history redundancy elimination. The performance of our model was evaluated in DVS128-Gesture dataset and other time-series datasets. The experimental results show that the proposed SRNN-SCBAM model makes better use of the history information in spatial and temporal dimensions with less memory space, and achieves higher accuracy compared to other models. Qi Xu 0008, Yuyuan Gao, Jiangrong Shen, Xuming Ran, Huajin Tang, Gang Pan 0001 |
NeurIPS | 1 |
| 2022 | Convolutional Neural Network Based Sleep Stage Classification with Class ImbalanceabstractAccurate sleep stage classification is vital to assess sleep quality and diagnose sleep disorders. Numerous deep learning based models have been designed for accomplishing this labor automatically. However, the class imbalance problem existing in polysomnography (PSG) datasets has been barely investigated in previous studies, which is one of the most challenging obstacles for the real-world sleep staging application. To address this issue, this paper proposes novel methods with signal-driven and image-driven ways of noise addition to balance the imbalanced relationship in the training dataset samples. We evaluate the effectiveness of the proposed methods which are integrated into a convolutional neural network (CNN) based model. Experimental results evaluated on Sleep-EDF-V1, Sleep-EDF and CCSHS databases demonstrate that the proposed balancing approaches with specific tensity Gaussian white noise could enhance the overall or stage N1 recognition to some degree, especially the combination of two types of Data augmentation (DA) strategies shows the superiority of overall accuracy improvement. Qi Xu 0008, Dongdong Zhou, Jian Wang 0112, Jiangrong Shen, Lauri Kettunen, Fengyu Cong |
IJCNN | 1 |
| 2022 | Detecting out-of-distribution samples via variational auto-encoder with reliable uncertainty estimation
Xuming Ran, Mingkun Xu, Lingrui Mei, Qi Xu 0008, Quanying Liu |
Neural Networks | 4 |
| 2022 | Robust Transcoding Sensory Information With Neural SpikesabstractNeural coding, including encoding and decoding, is one of the key problems in neuroscience for understanding how the brain uses neural signals to relate sensory perception and motor behaviors with neural systems. However, most of the existed studies only aim at dealing with the continuous signal of neural systems, while lacking a unique feature of biological neurons, termed spike, which is the fundamental information unit for neural computation as well as a building block for brain-machine interface. Aiming at these limitations, we propose a transcoding framework to encode multi-modal sensory information into neural spikes and then reconstruct stimuli from spikes. Sensory information can be compressed into 10% in terms of neural spikes, yet re-extract 100% of information by reconstruction. Our framework can not only feasibly and accurately reconstruct dynamical visual and auditory scenes, but also rebuild the stimulus patterns from functional magnetic resonance imaging (fMRI) brain activities. More importantly, it has a superb ability of noise immunity for various types of artificial noises and background signals. The proposed framework provides efficient ways to perform multimodal feature representation and reconstruction in a high-throughput fashion, with potential usage for efficient neuromorphic computing in a noisy environment. Qi Xu 0008, Jiangrong Shen, Xuming Ran, Huajin Tang, Gang Pan 0001, Jian K. Liu |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2020 | Deep CovDenseSNN: A hierarchical event-driven dynamic framework with spiking neurons in noisy environment
Qi Xu 0008, Jianxin Peng, Jiangrong Shen, Huajin Tang, Gang Pan 0001 |
Neural Networks | 1 |
| 2020 | Unsupervised AER Object Recognition Based on Multiscale Spatio-Temporal Features and Spiking NeuronsabstractThis article proposes an unsupervised address event representation (AER) object recognition approach. The proposed approach consists of a novel multiscale spatio-temporal feature (MuST) representation of input AER events and a spiking neural network (SNN) using spike-timing-dependent plasticity (STDP) for object recognition with MuST. MuST extracts the features contained in both the spatial and temporal information of AER event flow, and forms an informative and compact feature spike representation. We show not only how MuST exploits spikes to convey information more effectively, but also how it benefits the recognition using SNN. The recognition process is performed in an unsupervised manner, which does not need to specify the desired status of every single neuron of SNN, and thus can be flexibly applied in real-world recognition tasks. The experiments are performed on five AER datasets including a new one named GESTURE-DVS. Extensive experimental results show the effectiveness and advantages of the proposed approach. Qianhui Liu, Gang Pan 0001, Haibo Ruan, Dong Xing, Qi Xu 0008, Huajin Tang |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2019 | Overfitting remedy by sparsifying regularization on fully-connected layers of CNNs
Qi Xu 0008, Ming Zhang 0018, Zonghua Gu 0001, Gang Pan 0001 |
Neurocomputing | 1 |
| 2018 | CSNN: An Augmented Spiking based Framework with Perceptron-InceptionabstractSpiking Neural Networks (SNNs) represent and transmit information in spikes, which is considered more biologically realistic and computationally powerful than the traditional Artificial Neural Networks. The spiking neurons encode useful temporal information and possess highly anti-noise property. The feature extraction ability of typical SNNs is limited by shallow structures. This paper focuses on improving the feature extraction ability of SNNs in virtue of powerful feature extraction ability of Convolutional Neural Networks (CNNs). CNNs can extract abstract features resorting to the structure of the convolutional feature maps. We propose a CNN-SNN (CSNN) model to combine feature learning ability of CNNs with cognition ability of SNNs. The CSNN model learns the encoded spatial temporal representations of images in an event-driven way. We evaluate the CSNN model on the handwritten digits images dataset MNIST and its variational databases. In the presented experimental results, the proposed CSNN model is evaluated regarding learning capabilities, encoding mechanisms, robustness to noisy stimuli and its classification performance. The results show that CSNN behaves well compared to other cognitive models with significantly fewer neurons and training samples. Our work brings more biological realism into modern image classification models, with the hope that these models can inform how the brain performs this high-level vision task. Qi Xu 0008, Hang Yu 0010, Jiangrong Shen, Huajin Tang, Gang Pan 0001 |
IJCAI | 1 |
| 2017 | Darwin: A neuromorphic hardware co-processor based on spiking neural networks
De Ma, Juncheng Shen, Zonghua Gu 0001, Ming Zhang 0018, Xiaoqiang Xu, Qi Xu 0008, Yangjing Shen, Gang Pan 0001 |
J. Syst. Archit. | 7 |
| 2016 | Darwin: a neuromorphic hardware co-processor based on Spiking Neural Networks
Juncheng Shen, De Ma, Zonghua Gu 0001, Ming Zhang 0018, Xiaoqiang Xu, Qi Xu 0008, Yangjing Shen, Gang Pan 0001 |
Sci. China Inf. Sci. | 7 |