Xiongri Shen

dblp:319/8160 · DBLP profile ↗
← Back
4ranked-venue papers
1as first author
4since 2021 · last 2026
0000-0002-5655-6716ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-author · 3 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 2 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
2 papers
Deep learning architectures and training · 63% Speech recognition and synthesis · 20% Efficient and distributed learning · 17%
Interdisciplinary, comprehensive, and emerging computing
1 paper
Medical and health informatics · 100%
Computer architecture, parallel and distributed computing, and storage systems
1 paper
Emerging computing paradigms · 100%

Topics — the 7 heaviest of 7, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Deep learning architectures and training
spiking neural network
1.922026
SpikCommander: A High-performance Spiking Transformer with Multi-view Learning for Efficient Speech Command Recognition · AAAI 2026
S$^2$M-Former: Spiking Symmetric Mixing Branchformer for Brain Auditory Attention Detection · NeurIPS 2025
Natural language and speech › Speech recognition and synthesis › automatic speech recognition › keyword spotting
speech command recognition
1.012026
SpikCommander: A High-performance Spiking Transformer with Multi-view Learning for Efficient Speech Command Recognition · AAAI 2026
Machine learning › Deep learning architectures and training › spiking neural network
spiking transformer
1.012026
SpikCommander: A High-performance Spiking Transformer with Multi-view Learning for Efficient Speech Command Recognition · AAAI 2026
Machine learning › Efficient and distributed learning
energy-efficient learning
0.912025
S$^2$M-Former: Spiking Symmetric Mixing Branchformer for Brain Auditory Attention Detection · NeurIPS 2025
Medical and health informatics
brain-computer interface
0.912025
S$^2$M-Former: Spiking Symmetric Mixing Branchformer for Brain Auditory Attention Detection · NeurIPS 2025
Emerging computing paradigms
event-driven computing
0.312026
SpikCommander: A High-performance Spiking Transformer with Multi-view Learning for Efficient Speech Command Recognition · AAAI 2026
Machine learning › Deep learning architectures and training
transformer
0.312025
S$^2$M-Former: Spiking Symmetric Mixing Branchformer for Brain Auditory Attention Detection · NeurIPS 2025

Methods — techniques the papers use, named apart from their topics

spiking neural network · 3.7self-attention · 2.0multi-view learning · 2.0token-channel mixer · 1.7symmetric branch architecture · 1.7
YearPublicationVenuePosition
2026 SpikCommander: A High-performance Spiking Transformer with Multi-view Learning for Efficient Speech Command Recognition
abstract
Spiking neural networks (SNNs) offer a promising path toward energy-efficient speech command recognition (SCR) by leveraging their event-driven processing paradigm. However, existing SNN-based SCR methods often struggle to capture rich temporal dependencies and contextual information from speech due to limited temporal modeling and binary spike-based representations. To address these challenges, we first introduce the multi-view spiking temporal-aware self-attention (MSTASA) module, which combines effective spiking temporal-aware attention with a multi-view learning framework to model complementary temporal dependencies in speech commands. Building on MSTASA, we further propose SpikCommander, a fully spike-driven transformer architecture that integrates MSTASA with a spiking contextual refinement channel MLP (SCR-MLP) to jointly enhance temporal context modeling and channel-wise feature integration. We evaluate our method on three benchmark datasets: the Spiking Heidelberg Dataset (SHD), the Spiking Speech Commands (SSC), and the Google Speech Commands V2 (GSC). Extensive experiments demonstrate that SpikCommander consistently outperforms state-of-the-art (SOTA) SNN approaches with fewer parameters under comparable time steps, highlighting its effectiveness and efficiency for robust speech command recognition.
Jiaqi Wang 0003, Liutao Yu, Xiongri Shen, Sihang Guo, Chenlin Zhou, Zhiguo Zhang 0001, Zhengyu Ma
AAAI3
2025 Thread the Needle: Genomics-Guided Prompt-Bridged Attention Model for Survival Prediction of Glioma Based on MRI Images
Xubin Zheng, Xiongri Shen, Jiaqi Wang 0003, Zhenxi Song, Zhiguo Zhang 0001
MICCAI (7)3
2025 S$^2$M-Former: Spiking Symmetric Mixing Branchformer for Brain Auditory Attention Detection
abstract
Auditory attention detection (AAD) aims to decode listeners' focus in complex auditory environments from electroencephalography (EEG) recordings, which is crucial for developing neuro-steered hearing devices. Despite recent advancements, EEG-based AAD remains hindered by the absence of synergistic frameworks that can fully leverage complementary EEG features under energy-efficiency constraints. We propose ***S$^2$M-Former***, a novel ***s***piking ***s***ymmetric ***m***ixing framework to address this limitation through two key innovations: i) Presenting a spike-driven symmetric architecture composed of parallel spatial and frequency branches with mirrored modular design, leveraging biologically plausible token-channel mixers to enhance complementary learning across branches; ii) Introducing lightweight 1D token sequences to replace conventional 3D operations, reducing parameters by 14.7$\times$. The brain-inspired spiking architecture further reduces power consumption, achieving a 5.8$\times$ energy reduction compared to recent ANN methods, while also surpassing existing SNN baselines in terms of parameter efficiency and performance. Comprehensive experiments on three AAD benchmarks (KUL, DTU and AV-GC-AAD) across three settings (within-trial, cross-trial and cross-subject) demonstrate that S$^2$M-Former achieves comparable state-of-the-art (SOTA) decoding accuracy, making it a promising low-power, high-performance solution for AAD tasks. Code is available at https://github.com/JackieWang9811/S2M-Former.
Jiaqi Wang 0003, Zhengyu Ma, Xiongri Shen, Chenlin Zhou, Han Zhang 0035, Zhenxi Song, Zhiguo Zhang 0001
NeurIPS3
2024 GCAN: Generative Counterfactual Attention-Guided Network for Explainable Cognitive Decline Diagnostics Based on fMRI Functional Connectivity
Xiongri Shen, Zhenxi Song, Zhiguo Zhang 0001
MICCAI (10)1