VLDB 2026 Research / reviewers in the wild / expert
Shen Yan 0004
dblp:51/8939-4
· DBLP profile ↗
7ranked-venue papers
2as first author
7since 2021 · last 2025
0000-0002-8099-8960ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 5 · 2 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
3 papers |
Deep learning architectures and training · 84% Efficient and distributed learning · 10% Language models and text generation · 6% | |
| Computer architecture, parallel and distributed computing, and storage systems
2 papers |
Emerging computing paradigms · 90% Energy-efficient computing · 10% | |
| Theoretical computer science
1 paper |
Algorithms and data structures · 61% Information theory · 39% | |
| Databases, data mining, and information retrieval
1 paper |
Web and social media mining · 50% Data stream processing · 50% |
Topics — the 17 heaviest of 17, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Emerging computing paradigms
neuromorphic computing |
1.2 | 2 | 2023 | Towards Memory- and Time-Efficient Backpropagation for Training Spiking Neural Networks · ICCV 2023 Training High-Performance Low-Latency Spiking Neural Networks by Differentiation on Spike Representation · CVPR 2022 |
Machine learning › Deep learning architectures and training › mixture of experts
expert routing |
0.9 | 1 | 2025 | TC-MoE: Augmenting Mixture of Experts with Ternary Expert Choice · ICLR 2025 |
Machine learning › Deep learning architectures and training
mixture of experts |
0.9 | 1 | 2025 | TC-MoE: Augmenting Mixture of Experts with Ternary Expert Choice · ICLR 2025 |
Machine learning › Deep learning architectures and training › backpropagation
backpropagation through time |
0.7 | 1 | 2023 | Towards Memory- and Time-Efficient Backpropagation for Training Spiking Neural Networks · ICCV 2023 |
Machine learning › Deep learning architectures and training › spiking neural network
surrogate gradient |
0.7 | 1 | 2023 | Towards Memory- and Time-Efficient Backpropagation for Training Spiking Neural Networks · ICCV 2023 |
Emerging computing paradigms › neuromorphic computing
spiking neural network training |
0.7 | 1 | 2023 | Towards Memory- and Time-Efficient Backpropagation for Training Spiking Neural Networks · ICCV 2023 |
Machine learning › Deep learning architectures and training › spiking neural network
spiking neural network training |
0.6 | 1 | 2022 | Training High-Performance Low-Latency Spiking Neural Networks by Differentiation on Spike Representation · CVPR 2022 |
Algorithms and data structures › probabilistic data structures
bloom filter |
0.6 | 1 | 2022 | Elastic Bloom Filter: Deletable and Expandable Filter Using Elastic Fingerprints · IEEE Trans. Computers 2022 |
Information theory › hypothesis testing
false alarm probability |
0.6 | 1 | 2022 | Elastic Bloom Filter: Deletable and Expandable Filter Using Elastic Fingerprints · IEEE Trans. Computers 2022 |
Algorithms and data structures
probabilistic data structures |
0.6 | 1 | 2022 | Elastic Bloom Filter: Deletable and Expandable Filter Using Elastic Fingerprints · IEEE Trans. Computers 2022 |
Web and social media mining › event detection
burst detection |
0.5 | 1 | 2021 | BurstSketch: Finding Bursts in Data Streams · SIGMOD Conference 2021 |
Data stream processing
sketch |
0.5 | 1 | 2021 | BurstSketch: Finding Bursts in Data Streams · SIGMOD Conference 2021 |
Machine learning › Efficient and distributed learning › model compression › sparsity
activation sparsity |
0.3 | 1 | 2025 | TC-MoE: Augmenting Mixture of Experts with Ternary Expert Choice · ICLR 2025 |
Natural language and speech › Language models and text generation › efficient language model
large language model efficiency |
0.3 | 1 | 2025 | TC-MoE: Augmenting Mixture of Experts with Ternary Expert Choice · ICLR 2025 |
Energy-efficient computing
energy-efficient machine learning |
0.2 | 1 | 2023 | Towards Memory- and Time-Efficient Backpropagation for Training Spiking Neural Networks · ICCV 2023 |
Machine learning › Efficient and distributed learning › inference efficiency
low-latency inference |
0.2 | 1 | 2022 | Training High-Performance Low-Latency Spiking Neural Networks by Differentiation on Spike Representation · CVPR 2022 |
Information theory › signal processing › information embedding
fingerprints |
0.2 | 1 | 2022 | Elastic Bloom Filter: Deletable and Expandable Filter Using Elastic Fingerprints · IEEE Trans. Computers 2022 |
Methods — techniques the papers use, named apart from their topics
spatial learning through time · 1.3SLTT-K · 1.3surrogate gradient · 1.1firing rate coding · 1.1differentiation on spike representation · 1.1reward loss · 0.9load balance loss · 0.9elastic fingerprints · 0.6snapshotting · 0.5running track · 0.5
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | TC-MoE: Augmenting Mixture of Experts with Ternary Expert ChoiceabstractThe Mixture of Experts (MoE) architecture has emerged as a promising solution to reduce computational overhead by selectively activating subsets of model parameters.
The effectiveness of MoE models depends primarily on their routing mechanisms, with the widely adopted Top-K routing scheme used for activating experts.
However, the Top-K scheme has notable limitations,
including unnecessary activations and underutilization of experts.
In this work,
rather than modifying the routing mechanism as done in previous studies,
we propose the Ternary Choice MoE (TC-MoE),
a novel approach that expands the expert space by applying the ternary set {-1, 0, 1} to each expert.
This expansion allows more efficient and effective expert activations without incurring significant computational costs.
Additionally,
given the unique characteristics of the expanded expert space,
we introduce a new load balance loss and reward loss to ensure workload balance and achieve a flexible trade-off between effectiveness and efficiency.
Extensive experiments demonstrate that TC-MoE achieves an average improvement of over 1.1% compared with traditional approaches,
while reducing the average number of activated experts by up to 9%.
These results confirm that TC-MoE effectively addresses the inefficiencies of conventional routing schemes,
offering a more efficient and scalable solution for MoE-based large language models.
Code and models are available at https://github.com/stiger1000/TC-MoE. Shen Yan 0004, Xingyan Bin, Sijun Zhang, Yisen Wang 0001, Zhouchen Lin |
ICLR | 1 |
| 2024 | Sampling complex topology structures for spiking neural networks
Shen Yan 0004, Qingyan Meng, Mingqing Xiao 0002, Yisen Wang 0001, Zhouchen Lin |
Neural Networks | 1 |
| 2023 | Towards Memory- and Time-Efficient Backpropagation for Training Spiking Neural NetworksabstractSpiking Neural Networks (SNNs) are promising energy-efficient models for neuromorphic computing. For training the non-differentiable SNN methods, the backpropagation through time (BPTT) with surrogate gradients (SG) method has achieved high performance. However, this method suffers from considerable memory cost and training time during training. In this paper, we propose the Spatial Learning Through Time (SLTT) method that can achieve high performance while greatly improving training efficiency compared with BPTT. First, we show that the backpropagation of SNNs through the temporal domain contributes just a little to the final calculated gradients. Thus, we propose to ignore the unimportant routes in the computational graph during backpropagation. The proposed method reduces the number of scalar multiplications and achieves a small memory occupation that is independent of the total time steps. Furthermore, we propose a variant of SLTT, called SLTT-K, that allows backpropagation only at K time steps, then the required number of scalar multiplications is further reduced and is independent of the total time steps. Experiments on both static and neuromorphic datasets demonstrate superior training efficiency and performance of our SLTT. In particular, our method achieves state-of-the-art accuracy on ImageNet, while the memory cost and training time are reduced by more than 70% and 50%, respectively, compared with BPTT. Our code is available at https://github.com/qymeng94/SLTT. Qingyan Meng, Mingqing Xiao 0002, Shen Yan 0004, Yisen Wang 0001, Zhouchen Lin, Zhi-Quan Luo |
ICCV | 3 |
| 2022 | Training High-Performance Low-Latency Spiking Neural Networks by Differentiation on Spike RepresentationabstractSpiking Neural Network (SNN) is a promising energy-efficient AI model when implemented on neuromorphic hardware. However, it is a challenge to efficiently train SNNs due to their non-differentiability. Most existing methods either suffer from high latency (i.e., long simulation time steps), or cannot achieve as high performance as Artificial Neural Networks (ANNs). In this paper, we propose the Differentiation on Spike Representation (DSR) method, which could achieve high performance that is competitive to ANNs yet with low latency. First, we encode the spike trains into spike representation using (weighted) firing rate coding. Based on the spike representation, we systematically derive that the spiking dynamics with common neural models can be represented as some sub-differentiable mapping. With this viewpoint, our proposed DSR method trains SNNs through gradients of the mapping and avoids the common non-differentiability problem in SNN training. Then we analyze the error when representing the specific mapping with the forward computation of the SNN. To reduce such error, we propose to train the spike threshold in each layer, and to introduce a new hyperparameter for the neural models. With these components, the DSR method can achieve state-of-the-art SNN performance with low latency on both static and neuromorphic datasets, including CIFAR-10, CIFAR-100, ImageNet, and DVS-CIFAR10. Qingyan Meng, Mingqing Xiao 0002, Shen Yan 0004, Yisen Wang 0001, Zhouchen Lin, Zhi-Quan Luo |
CVPR | 3 |
| 2022 | Training much deeper spiking neural networks with a small number of time-steps
Qingyan Meng, Shen Yan 0004, Mingqing Xiao 0002, Yisen Wang 0001, Zhouchen Lin, Zhi-Quan Luo |
Neural Networks | 2 |
| 2022 | Elastic Bloom Filter: Deletable and Expandable Filter Using Elastic FingerprintsabstractThe Bloom filter, answering whether an item is in a set, has achieved great success in various fields, including networking, databases, and bioinformatics. However, the Bloom filter has two main shortcomings: no support of item deletion and no support of expansion. Existing solutions either support deletion at the cost of using additional memory, or support expansion at the cost of increasing the false positive rate and decreasing the query speed. Unlike existing solutions, we propose the Elastic Bloom filter (EBF) to address the two shortcomings simultaneously. Importantly, when EBF expands, the false positives decrease. Our key technique isElastic Fingerprints, which dynamically absorb and release bits during compression and expansion. To support deletion, EBF can first delete the corresponding fingerprint and then update the corresponding bit in the Bloom filter. To support expansion, Elastic Fingerprints release bits and insert them to the Bloom filter. Our experimental results show that the Elastic Bloom filter significantly outperforms existing works. Yuhan Wu 0001, Jintao He, Shen Yan 0004, Tong Yang 0003, Olivier Ruas, Gong Zhang 0001, Bin Cui 0001 |
IEEE Trans. Computers | 3 |
| 2021 | BurstSketch: Finding Bursts in Data StreamsabstractBurst is a common pattern in data streams which is characterized by a sudden increase in terms of arrival rate followed by a sudden decrease. Burst detection has attracted extensive attention from the research community. In this paper, we propose a novel sketch, namely BurstSketch, to detect bursts accurately in real time. BurstSketch first uses the technique Running Track to select potential burst items efficiently, and then monitors the potential burst items and capture the key features of burst pattern by a technique called Snapshotting. Experimental results show that our sketch achieves a 1.75 times higher recall rate than the strawman solution. Shen Yan 0004, Zikun Li, Decheng Tan, Tong Yang 0003, Bin Cui 0001 |
SIGMOD Conference | 2 |