Liutao Yu

dblp:311/9891 · DBLP profile ↗
← Back
7ranked-venue papers
0as first author
7since 2021 · last 2026
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 7 · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 3 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
6 papers
Deep learning architectures and training · 82% Speech recognition and synthesis · 14% Representation and self-supervised learning · 4%
Computer architecture, parallel and distributed computing, and storage systems
3 papers
Emerging computing paradigms · 94% Energy-efficient computing · 6%
Interdisciplinary, comprehensive, and emerging computing
3 papers
Bioinformatics and computational biology · 100%

Topics — the 14 heaviest of 14, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Bioinformatics and computational biology
computational neuroscience
2.332025
Time-Evolving Dynamical System for Learning Latent Representations of Mouse Visual Neural Activity · NeurIPS 2025
Long-Range Feedback Spiking Network Captures Dynamic and Static Representations of the Visual Cortex under Movie Stimuli · NeurIPS 2024
Deep Spiking Neural Networks with High Representation Similarity Model Visual Pathways of Macaque and Mouse · AAAI 2023
Machine learning › Deep learning architectures and training › spiking neural network
spiking transformer
2.022026
Spikingformer: A Key Foundation Model for Spiking Neural Networks · AAAI 2026
SpikCommander: A High-performance Spiking Transformer with Multi-view Learning for Efficient Speech Command Recognition · AAAI 2026
Machine learning › Deep learning architectures and training
transformer
1.822026
Spikingformer: A Key Foundation Model for Spiking Neural Networks · AAAI 2026
QKFormer: Hierarchical Spiking Transformer using Q-K Attention · NeurIPS 2024
Emerging computing paradigms
neuromorphic computing
1.822026
Spikingformer: A Key Foundation Model for Spiking Neural Networks · AAAI 2026
QKFormer: Hierarchical Spiking Transformer using Q-K Attention · NeurIPS 2024
Emerging computing paradigms › neuromorphic computing
spiking neural network
1.822026
Spikingformer: A Key Foundation Model for Spiking Neural Networks · AAAI 2026
QKFormer: Hierarchical Spiking Transformer using Q-K Attention · NeurIPS 2024
Machine learning › Deep learning architectures and training
spiking neural network
1.432026
SpikCommander: A High-performance Spiking Transformer with Multi-view Learning for Efficient Speech Command Recognition · AAAI 2026
Long-Range Feedback Spiking Network Captures Dynamic and Static Representations of the Visual Cortex under Movie Stimuli · NeurIPS 2024
Deep Spiking Neural Networks with High Representation Similarity Model Visual Pathways of Macaque and Mouse · AAAI 2023
Bioinformatics and computational biology › computational neuroscience
neural representation similarity
1.422024
Long-Range Feedback Spiking Network Captures Dynamic and Static Representations of the Visual Cortex under Movie Stimuli · NeurIPS 2024
Deep Spiking Neural Networks with High Representation Similarity Model Visual Pathways of Macaque and Mouse · AAAI 2023
Natural language and speech › Speech recognition and synthesis › automatic speech recognition › keyword spotting
speech command recognition
1.012026
SpikCommander: A High-performance Spiking Transformer with Multi-view Learning for Efficient Speech Command Recognition · AAAI 2026
Machine learning › Deep learning architectures and training
attention mechanism
0.812024
QKFormer: Hierarchical Spiking Transformer using Q-K Attention · NeurIPS 2024
Bioinformatics and computational biology › computational neuroscience › neural modeling
spiking neural network
0.812024
Long-Range Feedback Spiking Network Captures Dynamic and Static Representations of the Visual Cortex under Movie Stimuli · NeurIPS 2024
Emerging computing paradigms › neuromorphic computing › spiking neural network
spiking transformer
0.812024
QKFormer: Hierarchical Spiking Transformer using Q-K Attention · NeurIPS 2024
Energy-efficient computing
energy-efficient machine learning
0.312026
Spikingformer: A Key Foundation Model for Spiking Neural Networks · AAAI 2026
Emerging computing paradigms
event-driven computing
0.312026
SpikCommander: A High-performance Spiking Transformer with Multi-view Learning for Efficient Speech Command Recognition · AAAI 2026
Machine learning › Representation and self-supervised learning
contrastive learning
0.312025
Time-Evolving Dynamical System for Learning Latent Representations of Mouse Visual Neural Activity · NeurIPS 2025

Methods — techniques the papers use, named apart from their topics

self-attention · 4.0spiking neural network · 3.3representational similarity analysis · 2.8residual connections · 2.0multi-view learning · 2.0latent variable model · 1.7contrastive learning · 1.7spiking patch embedding · 1.5q-k attention · 1.5time-series analysis · 0.8time series analysis · 0.8vision transformer · 0.7convolutional neural network · 0.7
YearPublicationVenuePosition
2026 SpikCommander: A High-performance Spiking Transformer with Multi-view Learning for Efficient Speech Command Recognition
abstract
Spiking neural networks (SNNs) offer a promising path toward energy-efficient speech command recognition (SCR) by leveraging their event-driven processing paradigm. However, existing SNN-based SCR methods often struggle to capture rich temporal dependencies and contextual information from speech due to limited temporal modeling and binary spike-based representations. To address these challenges, we first introduce the multi-view spiking temporal-aware self-attention (MSTASA) module, which combines effective spiking temporal-aware attention with a multi-view learning framework to model complementary temporal dependencies in speech commands. Building on MSTASA, we further propose SpikCommander, a fully spike-driven transformer architecture that integrates MSTASA with a spiking contextual refinement channel MLP (SCR-MLP) to jointly enhance temporal context modeling and channel-wise feature integration. We evaluate our method on three benchmark datasets: the Spiking Heidelberg Dataset (SHD), the Spiking Speech Commands (SSC), and the Google Speech Commands V2 (GSC). Extensive experiments demonstrate that SpikCommander consistently outperforms state-of-the-art (SOTA) SNN approaches with fewer parameters under comparable time steps, highlighting its effectiveness and efficiency for robust speech command recognition.
Jiaqi Wang 0003, Liutao Yu, Xiongri Shen, Sihang Guo, Chenlin Zhou, Zhiguo Zhang 0001, Zhengyu Ma
AAAI2
2026 Spikingformer: A Key Foundation Model for Spiking Neural Networks
abstract
Spiking neural networks (SNNs) offer a promising energy-efficient alternative to artificial neural networks, due to their event-driven spiking computation. However, some foundation SNN backbones (including Spikformer and SEW ResNet) suffer from non-spike computations (integer-float multiplications) caused by the structure of their residual connections. These non-spike computations increase SNNs' power consumption and make them unsuitable for deployment on mainstream neuromorphic hardware. In this paper, we analyze the spike-driven behavior of the residual connection methods in SNNs. We then present Spikingformer, a novel spiking transformer backbone that merges the MS Residual connection with Self-Attention in a biologically plausible way to address the non-spike computation challenge in Spikformer while maintaining global modeling capabilities. We evaluate Spikingformer across 13 datasets spanning large static images, neuromorphic data, and natural language tasks, and demonstrate the effectiveness and universality of Spikingformer, setting a vital benchmark for spiking neural networks. In addition, with the spike-driven features and global modeling capabilities, Spikingformer is expected to become a more efficient general-purpose SNN backbone towards energy-efficient artificial intelligence.
Chenlin Zhou, Liutao Yu, Zhaokun Zhou, Han Zhang 0035, Jiaqi Wang 0003, Zhengyu Ma, Yonghong Tian 0001
AAAI2
2026 Efficient speech command recognition leveraging spiking neural networks and progressive time-scaled curriculum distillation
Jiaqi Wang 0003, Liutao Yu, Liwei Huang, Chenlin Zhou, Han Zhang 0035, Zhenxi Song, Honghai Liu 0001, Min Zhang 0005, Zhengyu Ma, Zhiguo Zhang 0001
Neural Networks2
2025 Time-Evolving Dynamical System for Learning Latent Representations of Mouse Visual Neural Activity
abstract
Seeking high-quality representations with latent variable models (LVMs) to reveal the intrinsic correlation between neural activity and behavior or sensory stimuli has attracted much interest. In the study of the biological visual system, naturalistic visual stimuli are inherently high-dimensional and time-dependent, leading to intricate dynamics within visual neural activity. However, most work on LVMs has not explicitly considered neural temporal relationships. To cope with such conditions, we propose Time-Evolving Visual Dynamical System (TE-ViDS), a sequential LVM that decomposes neural activity into low-dimensional latent representations that evolve over time. To better align the model with the characteristics of visual neural activity, we split latent representations into two parts and apply contrastive learning to shape them. Extensive experiments on synthetic datasets and real neural datasets from the mouse visual cortex demonstrate that TE-ViDS achieves the best decoding performance on naturalistic scenes/movies, extracts interpretable latent trajectories that uncover clear underlying neural dynamics, and provides new insights into differences in visual information processing between subjects and between cortical regions. In summary, TE-ViDS is markedly competent in extracting stimulus-relevant embeddings from visual neural activity and contributes to the understanding of visual processing mechanisms. Our codes are available at https://github.com/Grasshlw/Time-Evolving-Visual-Dynamical-System.
Liwei Huang, Zhengyu Ma, Liutao Yu, Yonghong Tian 0001
NeurIPS3
2024 Long-Range Feedback Spiking Network Captures Dynamic and Static Representations of the Visual Cortex under Movie Stimuli
abstract
Deep neural networks (DNNs) are widely used models for investigating biological visual representations. However, existing DNNs are mostly designed to analyze neural responses to static images, relying on feedforward structures and lacking physiological neuronal mechanisms. There is limited insight into how the visual cortex represents natural movie stimuli that contain context-rich information. To address these problems, this work proposes the long-range feedback spiking network (LoRaFB-SNet), which mimics top-down connections between cortical regions and incorporates spike information processing mechanisms inherent to biological neurons. Taking into account the temporal dependence of representations under movie stimuli, we present Time-Series Representational Similarity Analysis (TSRSA) to measure the similarity between model representations and visual cortical representations of mice. LoRaFB-SNet exhibits the highest level of representational similarity, outperforming other well-known and leading alternatives across various experimental paradigms, especially when representing long movie stimuli. We further conduct experiments to quantify how temporal structures (dynamic information) and static textures (static information) of the movie stimuli influence representational similarity, suggesting that our model benefits from long-range feedback to encode context-dependent representations just like the brain. Altogether, LoRaFB-SNet is highly competent in capturing both dynamic and static representations of the mouse visual cortex and contributes to the understanding of movie processing mechanisms of the visual system. Our codes are available at https://github.com/Grasshlw/SNN-Neural-Similarity-Movie.
Liwei Huang, Zhengyu Ma, Liutao Yu, Yonghong Tian 0001
NeurIPS3
2024 QKFormer: Hierarchical Spiking Transformer using Q-K Attention
abstract
Spiking Transformers, which integrate Spiking Neural Networks (SNNs) with Transformer architectures, have attracted significant attention due to their potential for low energy consumption and high performance. However, there remains a substantial gap in performance between SNNs and Artificial Neural Networks (ANNs). To narrow this gap, we have developed QKFormer, a direct training spiking transformer with the following features: i) _Linear complexity and high energy efficiency_, the novel spike-form Q-K attention module efficiently models the token or channel attention through binary vectors and enables the construction of larger models. ii) _Multi-scale spiking representation_, achieved by a hierarchical structure with the different numbers of tokens across blocks. iii) _Spiking Patch Embedding with Deformed Shortcut (SPEDS)_, enhances spiking information transmission and integration, thus improving overall performance. It is shown that QKFormer achieves significantly superior performance over existing state-of-the-art SNN models on various mainstream datasets. Notably, with comparable size to Spikformer (66.34 M, 74.81\%), QKFormer (64.96 M) achieves a groundbreaking top-1 accuracy of **85.65\%** on ImageNet-1k, substantially outperforming Spikformer by **10.84\%**. To our best knowledge, this is the first time that directly training SNNs have exceeded 85\% accuracy on ImageNet-1K.
Chenlin Zhou, Han Zhang 0035, Zhaokun Zhou, Liutao Yu, Liwei Huang, Xiaopeng Fan 0001, Li Yuan 0007, Zhengyu Ma, Yonghong Tian 0001
NeurIPS4
2023 Deep Spiking Neural Networks with High Representation Similarity Model Visual Pathways of Macaque and Mouse
abstract
Deep artificial neural networks (ANNs) play a major role in modeling the visual pathways of primate and rodent. However, they highly simplify the computational properties of neurons compared to their biological counterparts. Instead, Spiking Neural Networks (SNNs) are more biologically plausible models since spiking neurons encode information with time sequences of spikes, just like biological neurons do. However, there is a lack of studies on visual pathways with deep SNNs models. In this study, we model the visual cortex with deep SNNs for the first time, and also with a wide range of state-of-the-art deep CNNs and ViTs for comparison. Using three similarity metrics, we conduct neural representation similarity experiments on three neural datasets collected from two species under three types of stimuli. Based on extensive similarity analyses, we further investigate the functional hierarchy and mechanisms across species. Almost all similarity scores of SNNs are higher than their counterparts of CNNs with an average of 6.6%. Depths of the layers with the highest similarity scores exhibit little differences across mouse cortical regions, but vary significantly across macaque regions, suggesting that the visual processing structure of mice is more regionally homogeneous than that of macaques. Besides, the multi-branch structures observed in some top mouse brain-like neural networks provide computational evidence of parallel processing streams in mice, and the different performance in fitting macaque neural representations under different stimuli exhibits the functional specialization of information processing in macaques. Taken together, our study demonstrates that SNNs could serve as promising candidates to better model and explain the functional hierarchy and mechanisms of the visual system.
Liwei Huang, Zhengyu Ma, Liutao Yu, Yonghong Tian 0001
AAAI3