VLDB 2026 Research / reviewers in the wild / expert
Yanqi Chen
dblp:284/9379
· DBLP profile ↗
13ranked-venue papers
7as first author
13since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 8 · 3 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-author · 3 since 2021Security and privacy · 2 · 2 first-author · 2 since 2021Systems, architecture and hardware · 1 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | HASE: Hardware-Aware Scheduling for Inference Tasks in Heterogeneous GPU ClustersabstractWhen processing large-scale co-located inference workloads in heterogeneous GPU clusters, existing cluster scheduling mechanisms often increase the job makespan due to neglecting the mutual performance interference and resource contention for co-located tasks. This deficiency is dramatically amplified in typical neural networks based deep learning workloads because these workloads are highly sensitive to hardware configurations like GPU memory capacity, bandwidth and GPU core frequency. Therefore, a hardware-performance-aware scheduler capable of adaptively and dynamically dispatching tasks according to GPU hardware characteristics is crucial for reducing the overall completion time of job queues. To address this issue, we propose Hardware-Aware Scheduling for Inference Tasks in Heterogeneous GPU Clusters (HASE), a novel scheduling framework that dynamically adapts to GPU hardware characteristics for real-time optimal task placement.HASEconsists of two core components: a kernel-level latency prediction model and a hybrid scheduling strategy. Unlike conventional approaches that rely on coarse-grained model-level features, our predictor decomposes inference models into fine-grained computational kernels using ONNX Runtime graph optimization, and predicts individual kernel execution times under varying GPU load conditions. By integrating static hardware specifications, dynamic microbenchmarks, and real-time DCGM profiling metrics, the predictor captures both operator-level heterogeneity and background load interference. The hybrid scheduling strategy combines a two-stage greedy search for task placement with a resource reservation and backfilling mechanism to balance immediate optimization and long-term fairness. Experimental results on CNN-based inference workloads (YOLO, ResNet, VGG, DenseNet, and MobileNet series) demonstrate thatHASEachieves a kernel-level prediction accuracy of 8.2% MAPE and model-level accuracy of$R^{2}$= 0.91. Moreover,HASEreduces 51% total job makespan compared to traditional round-robin scheduling, and maintains per-task scheduling decision time under one second in clusters with up to hundreds of GPUs. Yanqi Chen, Congfeng Jiang, Chunpeng Wu, Qinghe Ye, Jianing Niu, Lingjia Lao |
IEEE Trans. Cloud Comput. | 1 |
| 2025 | Efficient Reconstruction of S-Boxes from Partial Cryptographic Tables via MILP Modeling
Yanqi Chen |
Inscrypt (1) | 1 |
| 2025 | DiskJoin: Large-scale Vector Similarity Join with SSDabstractSimilarity join-a widely used operation in data science-finds all pairs of items that have distance smaller than a threshold. Prior work has explored distributed computation methods to scale similarity join to large data volumes but these methods require a cluster deployment, and efficiency suffers from expensive inter-machine communication. On the other hand, disk-based solutions are more cost-effective by using a single machine and storing the large dataset on high-performance external storage, such as NVMe SSDs, but in these methods the disk I/O time is a serious bottleneck. In this paper, we propose DiskJoin, the first disk-based similarity join algorithm that can process billion-scale vector datasets efficiently on a single machine. DiskJoin improves disk I/O by tailoring the data access patterns to avoid repetitive accesses and read amplification. It also uses main memory as a dynamic cache and carefully manages cache eviction to improve cache hit rate and reduce disk retrieval time. For further acceleration, we adopt a probabilistic pruning technique that can effectively prune a large number of vector pairs from computation. Our evaluation on real-world, large-scale datasets shows that DiskJoin significantly outperforms alternatives, achieving speedups from 50× to 1000×. Yanqi Chen, Xiao Yan 0002, Alexandra Meliou, Eric Lo 0001 |
Proc. ACM Manag. Data | 1 |
| 2024 | Cube Attacks Against Trivium, Kreyvium and ACORN with Practical Complexity
Yanqi Chen |
Inscrypt (2) | 1 |
| 2024 | Temporal Contrastive Learning for Spiking Neural Networks
Haonan Qiu, Zeyin Song, Yanqi Chen, Munan Ning, Wei Fang 0006, Zhengyu Ma, Li Yuan 0007, Yonghong Tian 0001 |
ICANN (10) | 3 |
| 2024 | Self-architectural knowledge distillation for spiking neural networks
Haonan Qiu, Munan Ning, Zeyin Song, Wei Fang 0006, Yanqi Chen, Zhengyu Ma, Li Yuan 0007, Yonghong Tian 0001 |
Neural Networks | 5 |
| 2023 | A Unified Framework for Soft Threshold Pruning
Yanqi Chen, Zhengyu Ma, Wei Fang 0006, Xiawu Zheng, Zhaofei Yu, Yonghong Tian 0001 |
ICLR | 1 |
| 2023 | Parallel Spiking Neurons with High Efficiency and Ability to Learn Long-term DependenciesabstractVanilla spiking neurons in Spiking Neural Networks (SNNs) use charge-fire-reset neuronal dynamics, which can only be simulated serially and can hardly learn long-time dependencies. We find that when removing reset, the neuronal dynamics can be reformulated in a non-iterative form and parallelized. By rewriting neuronal dynamics without reset to a general formulation, we propose the Parallel Spiking Neuron (PSN), which generates hidden states that are independent of their predecessors, resulting in parallelizable neuronal dynamics and extremely high simulation speed. The weights of inputs in the PSN are fully connected, which maximizes the utilization of temporal information. To avoid the use of future inputs for step-by-step inference, the weights of the PSN can be masked, resulting in the masked PSN. By sharing weights across time-steps based on the masked PSN, the sliding PSN is proposed to handle sequences of varying lengths. We evaluate the PSN family on simulation speed and temporal/static data classification, and the results show the overwhelming advantage of the PSN family in efficiency and accuracy. To the best of our knowledge, this is the first study about parallelizing spiking neurons and can be a cornerstone for the spiking deep learning research. Our codes are available at https://github.com/fangwei123456/Parallel-Spiking-Neuron. Wei Fang 0006, Zhaofei Yu, Zhaokun Zhou, Yanqi Chen, Zhengyu Ma, Timothée Masquelier, Yonghong Tian 0001 |
NeurIPS | 5 |
| 2022 | State Transition of Dendritic Spines Improves Learning of Sparse Spiking Neural NetworksabstractSpiking Neural Networks (SNNs) are considered a promising alternative to Artificial Neural Networks (ANNs) for their event-driven computing paradigm when deployed on energy-efficient neuromorphic hardware. Recently, deep SNNs have shown breathtaking performance improvement through cutting-edge training strategy and flexible structure, which also scales up the number of parameters and computational burdens in a single network. Inspired by the state transition of dendritic spines in the filopodial model of spinogenesis, we model different states of SNN weights, facilitating weight optimization for pruning. Furthermore, the pruning speed can be regulated by using different functions describing the growing threshold of state transition. We organize these techniques as a dynamic pruning algorithm based on nonlinear reparameterization mapping from spine size to SNN weights. Our approach yields sparse deep networks on the large-scale dataset (SEW ResNet18 on ImageNet) while maintaining state-of-the-art low performance loss ( 3% at 88.8% sparsity) compared to existing pruning methods on directly trained SNNs. Moreover, we find out pruning speed regulation while learning is crucial to avoiding disastrous performance degradation at the final stages of training, which may shed light on future work on SNN pruning. Yanqi Chen, Zhaofei Yu, Wei Fang 0006, Zhengyu Ma, Tiejun Huang 0001, Yonghong Tian 0001 |
ICML | 1 |
| 2022 | On-line harmonic signal denoising from the measurement with non-stationary and non-Gaussian noise
Liang Yu 0003, Yanqi Chen, Yongli Zhang, Ran Wang 0011, Zhaodong Zhang |
Signal Process. | 2 |
| 2021 | Incorporating Learnable Membrane Time Constant to Enhance Learning of Spiking Neural NetworksabstractSpiking Neural Networks (SNNs) have attracted enormous research interest due to temporal information processing capability, low power consumption, and high biological plausibility. However, the formulation of efficient and high-performance learning algorithms for SNNs is still challenging. Most existing learning methods learn weights only, and require manual tuning of the membrane-related parameters that determine the dynamics of a single spiking neuron. These parameters are typically chosen to be the same for all neurons, which limits the diversity of neurons and thus the expressiveness of the resulting SNNs. In this paper, we take inspiration from the observation that membrane-related parameters are different across brain regions, and propose a training algorithm that is capable of learning not only the synaptic weights but also the membrane time constants of SNNs. We show that incorporating learnable membrane time constants can make the network less sensitive to initial values and can speed up learning. In addition, we reevaluate the pooling methods in SNNs and find that max-pooling will not lead to significant information loss and have the advantage of low computation cost and binary compatibility. We evaluate the proposed method for image classification tasks on both traditional static MNIST, Fashion-MNIST, CIFAR-10 datasets, and neuromorphic N-MNIST, CIFAR10-DVS, DVS128 Gesture datasets. The experiment results show that the proposed method outperforms the state-of-the-art accuracy on nearly all datasets, using fewer time-steps. Our codes are available at https://github.com/fangwei123456/Parametric-Leaky-Integrate-and-Fire-Spiking-Neuron. Wei Fang 0006, Zhaofei Yu, Yanqi Chen, Timothée Masquelier, Tiejun Huang 0001, Yonghong Tian 0001 |
ICCV | 3 |
| 2021 | Pruning of Deep Spiking Neural Networks through Gradient RewiringabstractSpiking Neural Networks (SNNs) have been attached great importance due to their biological plausibility and high energy-efficiency on neuromorphic chips. As these chips are usually resource-constrained, the compression of SNNs is thus crucial along the road of practical use of SNNs. Most existing methods directly apply pruning approaches in artificial neural networks (ANNs) to SNNs, which ignore the difference between ANNs and SNNs, thus limiting the performance of the pruned SNNs. Besides, these methods are only suitable for shallow SNNs. In this paper, inspired by synaptogenesis and synapse elimination in the neural system, we propose gradient rewiring (Grad R), a joint learning algorithm of connectivity and weight for SNNs, that enables us to seamlessly optimize network structure without retraining. Our key innovation is to redefine the gradient to a new synaptic parameter, allowing better exploration of network structures by taking full advantage of the competition between pruning and regrowth of connections. The experimental results show that the proposed method achieves minimal loss of SNNs' performance on MNIST and CIFAR-10 datasets so far. Moreover, it reaches a ~3.5% accuracy loss under unprecedented 0.73% connectivity, which reveals remarkable structure refining capability in SNNs. Our work suggests that there exists extremely high redundancy in deep SNNs. Our codes are available at https://github.com/Yanqi-Chen/Gradient-Rewiring. Yanqi Chen, Zhaofei Yu, Wei Fang 0006, Tiejun Huang 0001, Yonghong Tian 0001 |
IJCAI | 1 |
| 2021 | Deep Residual Learning in Spiking Neural NetworksabstractDeep Spiking Neural Networks (SNNs) present optimization difficulties for gradient-based approaches due to discrete binary activation and complex spatial-temporal dynamics. Considering the huge success of ResNet in deep learning, it would be natural to train deep SNNs with residual learning. Previous Spiking ResNet mimics the standard residual block in ANNs and simply replaces ReLU activation layers with spiking neurons, which suffers the degradation problem and can hardly implement residual learning. In this paper, we propose the spike-element-wise (SEW) ResNet to realize residual learning in deep SNNs. We prove that the SEW ResNet can easily implement identity mapping and overcome the vanishing/exploding gradient problems of Spiking ResNet. We evaluate our SEW ResNet on ImageNet, DVS Gesture, and CIFAR10-DVS datasets, and show that SEW ResNet outperforms the state-of-the-art directly trained SNNs in both accuracy and time-steps. Moreover, SEW ResNet can achieve higher performance by simply adding more layers, providing a simple method to train deep SNNs. To our best knowledge, this is the first time that directly training deep SNNs with more than 100 layers becomes possible. Our codes are available at https://github.com/fangwei123456/Spike-Element-Wise-ResNet. Wei Fang 0006, Zhaofei Yu, Yanqi Chen, Tiejun Huang 0001, Timothée Masquelier, Yonghong Tian 0001 |
NeurIPS | 3 |