Shen Yan 0004

dblp:51/8939-4 · DBLP profile ↗
← Back
7ranked-venue papers
2as first author
7since 2021 · last 2025
0000-0002-8099-8960ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 5 · 2 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
3 papers
Deep learning architectures and training · 84% Efficient and distributed learning · 10% Language models and text generation · 6%
Computer architecture, parallel and distributed computing, and storage systems
2 papers
Emerging computing paradigms · 90% Energy-efficient computing · 10%
Theoretical computer science
1 paper
Algorithms and data structures · 61% Information theory · 39%
Databases, data mining, and information retrieval
1 paper
Web and social media mining · 50% Data stream processing · 50%

Topics — the 17 heaviest of 17, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Emerging computing paradigms
neuromorphic computing
1.222023
Towards Memory- and Time-Efficient Backpropagation for Training Spiking Neural Networks · ICCV 2023
Training High-Performance Low-Latency Spiking Neural Networks by Differentiation on Spike Representation · CVPR 2022
Machine learning › Deep learning architectures and training › mixture of experts
expert routing
0.912025
TC-MoE: Augmenting Mixture of Experts with Ternary Expert Choice · ICLR 2025
Machine learning › Deep learning architectures and training
mixture of experts
0.912025
TC-MoE: Augmenting Mixture of Experts with Ternary Expert Choice · ICLR 2025
Machine learning › Deep learning architectures and training › backpropagation
backpropagation through time
0.712023
Towards Memory- and Time-Efficient Backpropagation for Training Spiking Neural Networks · ICCV 2023
Machine learning › Deep learning architectures and training › spiking neural network
surrogate gradient
0.712023
Towards Memory- and Time-Efficient Backpropagation for Training Spiking Neural Networks · ICCV 2023
Emerging computing paradigms › neuromorphic computing
spiking neural network training
0.712023
Towards Memory- and Time-Efficient Backpropagation for Training Spiking Neural Networks · ICCV 2023
Machine learning › Deep learning architectures and training › spiking neural network
spiking neural network training
0.612022
Training High-Performance Low-Latency Spiking Neural Networks by Differentiation on Spike Representation · CVPR 2022
Algorithms and data structures › probabilistic data structures
bloom filter
0.612022
Elastic Bloom Filter: Deletable and Expandable Filter Using Elastic Fingerprints · IEEE Trans. Computers 2022
Information theory › hypothesis testing
false alarm probability
0.612022
Elastic Bloom Filter: Deletable and Expandable Filter Using Elastic Fingerprints · IEEE Trans. Computers 2022
Algorithms and data structures
probabilistic data structures
0.612022
Elastic Bloom Filter: Deletable and Expandable Filter Using Elastic Fingerprints · IEEE Trans. Computers 2022
Web and social media mining › event detection
burst detection
0.512021
BurstSketch: Finding Bursts in Data Streams · SIGMOD Conference 2021
Data stream processing
sketch
0.512021
BurstSketch: Finding Bursts in Data Streams · SIGMOD Conference 2021
Machine learning › Efficient and distributed learning › model compression › sparsity
activation sparsity
0.312025
TC-MoE: Augmenting Mixture of Experts with Ternary Expert Choice · ICLR 2025
Natural language and speech › Language models and text generation › efficient language model
large language model efficiency
0.312025
TC-MoE: Augmenting Mixture of Experts with Ternary Expert Choice · ICLR 2025
Energy-efficient computing
energy-efficient machine learning
0.212023
Towards Memory- and Time-Efficient Backpropagation for Training Spiking Neural Networks · ICCV 2023
Machine learning › Efficient and distributed learning › inference efficiency
low-latency inference
0.212022
Training High-Performance Low-Latency Spiking Neural Networks by Differentiation on Spike Representation · CVPR 2022
Information theory › signal processing › information embedding
fingerprints
0.212022
Elastic Bloom Filter: Deletable and Expandable Filter Using Elastic Fingerprints · IEEE Trans. Computers 2022

Methods — techniques the papers use, named apart from their topics

spatial learning through time · 1.3SLTT-K · 1.3surrogate gradient · 1.1firing rate coding · 1.1differentiation on spike representation · 1.1reward loss · 0.9load balance loss · 0.9elastic fingerprints · 0.6snapshotting · 0.5running track · 0.5
YearPublicationVenuePosition
2025 TC-MoE: Augmenting Mixture of Experts with Ternary Expert Choice
abstract
The Mixture of Experts (MoE) architecture has emerged as a promising solution to reduce computational overhead by selectively activating subsets of model parameters. The effectiveness of MoE models depends primarily on their routing mechanisms, with the widely adopted Top-K routing scheme used for activating experts. However, the Top-K scheme has notable limitations, including unnecessary activations and underutilization of experts. In this work, rather than modifying the routing mechanism as done in previous studies, we propose the Ternary Choice MoE (TC-MoE), a novel approach that expands the expert space by applying the ternary set {-1, 0, 1} to each expert. This expansion allows more efficient and effective expert activations without incurring significant computational costs. Additionally, given the unique characteristics of the expanded expert space, we introduce a new load balance loss and reward loss to ensure workload balance and achieve a flexible trade-off between effectiveness and efficiency. Extensive experiments demonstrate that TC-MoE achieves an average improvement of over 1.1% compared with traditional approaches, while reducing the average number of activated experts by up to 9%. These results confirm that TC-MoE effectively addresses the inefficiencies of conventional routing schemes, offering a more efficient and scalable solution for MoE-based large language models. Code and models are available at https://github.com/stiger1000/TC-MoE.
Shen Yan 0004, Xingyan Bin, Sijun Zhang, Yisen Wang 0001, Zhouchen Lin
ICLR1
2024 Sampling complex topology structures for spiking neural networks
Shen Yan 0004, Qingyan Meng, Mingqing Xiao 0002, Yisen Wang 0001, Zhouchen Lin
Neural Networks1
2023 Towards Memory- and Time-Efficient Backpropagation for Training Spiking Neural Networks
abstract
Spiking Neural Networks (SNNs) are promising energy-efficient models for neuromorphic computing. For training the non-differentiable SNN methods, the backpropagation through time (BPTT) with surrogate gradients (SG) method has achieved high performance. However, this method suffers from considerable memory cost and training time during training. In this paper, we propose the Spatial Learning Through Time (SLTT) method that can achieve high performance while greatly improving training efficiency compared with BPTT. First, we show that the backpropagation of SNNs through the temporal domain contributes just a little to the final calculated gradients. Thus, we propose to ignore the unimportant routes in the computational graph during backpropagation. The proposed method reduces the number of scalar multiplications and achieves a small memory occupation that is independent of the total time steps. Furthermore, we propose a variant of SLTT, called SLTT-K, that allows backpropagation only at K time steps, then the required number of scalar multiplications is further reduced and is independent of the total time steps. Experiments on both static and neuromorphic datasets demonstrate superior training efficiency and performance of our SLTT. In particular, our method achieves state-of-the-art accuracy on ImageNet, while the memory cost and training time are reduced by more than 70% and 50%, respectively, compared with BPTT. Our code is available at https://github.com/qymeng94/SLTT.
Qingyan Meng, Mingqing Xiao 0002, Shen Yan 0004, Yisen Wang 0001, Zhouchen Lin, Zhi-Quan Luo
ICCV3
2022 Training High-Performance Low-Latency Spiking Neural Networks by Differentiation on Spike Representation
abstract
Spiking Neural Network (SNN) is a promising energy-efficient AI model when implemented on neuromorphic hardware. However, it is a challenge to efficiently train SNNs due to their non-differentiability. Most existing methods either suffer from high latency (i.e., long simulation time steps), or cannot achieve as high performance as Artificial Neural Networks (ANNs). In this paper, we propose the Differentiation on Spike Representation (DSR) method, which could achieve high performance that is competitive to ANNs yet with low latency. First, we encode the spike trains into spike representation using (weighted) firing rate coding. Based on the spike representation, we systematically derive that the spiking dynamics with common neural models can be represented as some sub-differentiable mapping. With this viewpoint, our proposed DSR method trains SNNs through gradients of the mapping and avoids the common non-differentiability problem in SNN training. Then we analyze the error when representing the specific mapping with the forward computation of the SNN. To reduce such error, we propose to train the spike threshold in each layer, and to introduce a new hyperparameter for the neural models. With these components, the DSR method can achieve state-of-the-art SNN performance with low latency on both static and neuromorphic datasets, including CIFAR-10, CIFAR-100, ImageNet, and DVS-CIFAR10.
Qingyan Meng, Mingqing Xiao 0002, Shen Yan 0004, Yisen Wang 0001, Zhouchen Lin, Zhi-Quan Luo
CVPR3
2022 Training much deeper spiking neural networks with a small number of time-steps
Qingyan Meng, Shen Yan 0004, Mingqing Xiao 0002, Yisen Wang 0001, Zhouchen Lin, Zhi-Quan Luo
Neural Networks2
2022 Elastic Bloom Filter: Deletable and Expandable Filter Using Elastic Fingerprints
abstract
The Bloom filter, answering whether an item is in a set, has achieved great success in various fields, including networking, databases, and bioinformatics. However, the Bloom filter has two main shortcomings: no support of item deletion and no support of expansion. Existing solutions either support deletion at the cost of using additional memory, or support expansion at the cost of increasing the false positive rate and decreasing the query speed. Unlike existing solutions, we propose the Elastic Bloom filter (EBF) to address the two shortcomings simultaneously. Importantly, when EBF expands, the false positives decrease. Our key technique isElastic Fingerprints, which dynamically absorb and release bits during compression and expansion. To support deletion, EBF can first delete the corresponding fingerprint and then update the corresponding bit in the Bloom filter. To support expansion, Elastic Fingerprints release bits and insert them to the Bloom filter. Our experimental results show that the Elastic Bloom filter significantly outperforms existing works.
Yuhan Wu 0001, Jintao He, Shen Yan 0004, Tong Yang 0003, Olivier Ruas, Gong Zhang 0001, Bin Cui 0001
IEEE Trans. Computers3
2021 BurstSketch: Finding Bursts in Data Streams
abstract
Burst is a common pattern in data streams which is characterized by a sudden increase in terms of arrival rate followed by a sudden decrease. Burst detection has attracted extensive attention from the research community. In this paper, we propose a novel sketch, namely BurstSketch, to detect bursts accurately in real time. BurstSketch first uses the technique Running Track to select potential burst items efficiently, and then monitors the potential burst items and capture the key features of burst pattern by a technique called Snapshotting. Experimental results show that our sketch achieves a 1.75 times higher recall rate than the strawman solution.
Shen Yan 0004, Zikun Li, Decheng Tan, Tong Yang 0003, Bin Cui 0001
SIGMOD Conference2