Shiye Lei

dblp:243/3399 · DBLP profile ↗
← Back
8ranked-venue papers
6as first author
8since 2021 · last 2026
0000-0001-7810-9346ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 7 · 6 first-author · 7 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
5 papers
Reinforcement learning · 37% Efficient and distributed learning · 35% Deep learning architectures and training · 13%

Topics — the 13 heaviest of 15, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Efficient and distributed learning
data-efficient learning
1.222026
Offline Behavioral Data Selection · KDD (1) 2026
A Comprehensive Survey of Dataset Distillation · IEEE Trans. Pattern Anal. Mach. Intell. 2024
Machine learning › Reinforcement learning › imitation learning › offline imitation learning
behavior cloning
1.012026
Offline Behavioral Data Selection · KDD (1) 2026
Machine learning › Efficient and distributed learning
data selection
1.012026
Offline Behavioral Data Selection · KDD (1) 2026
Machine learning › Efficient and distributed learning › data selection
data subset selection
1.012026
Offline Behavioral Data Selection · KDD (1) 2026
Machine learning › Reinforcement learning › offline reinforcement learning
offline policy learning
1.012026
Offline Behavioral Data Selection · KDD (1) 2026
Computer vision › Vision and language
image captioning
0.912025
Image Captions are Natural Prompts for Training Data Synthesis · Int. J. Comput. Vis. 2025
Machine learning › Deep learning architectures and training › data-centric deep learning
training data generation
0.912025
Image Captions are Natural Prompts for Training Data Synthesis · Int. J. Comput. Vis. 2025
Machine learning › Efficient and distributed learning
dataset distillation
0.812024
A Comprehensive Survey of Dataset Distillation · IEEE Trans. Pattern Anal. Mach. Intell. 2024
Machine learning › Reinforcement learning
offline reinforcement learning
0.812024
Offline Behavior Distillation · NeurIPS 2024
Machine learning › Reinforcement learning
policy learning
0.812024
Offline Behavior Distillation · NeurIPS 2024
Machine learning › Reinforcement learning › sample efficiency
sample-efficient policy learning
0.812024
Offline Behavior Distillation · NeurIPS 2024
Machine learning › Deep learning architectures and training
complex-valued neural network
0.712023
Spectral complexity-scaled generalisation bound of complex-valued neural networks · Artif. Intell. 2023
Machine learning › Learning theory
generalization bounds
0.712023
Spectral complexity-scaled generalisation bound of complex-valued neural networks · Artif. Intell. 2023

Methods — techniques the papers use, named apart from their topics

dual ranking · 1.0behavioral cloning · 1.0language model · 0.9meta-learning · 0.8data matching · 0.8bi-level optimization · 0.8behavior distillation · 0.8spectral norm · 0.7maurey sparsification lemma · 0.7dudley entropy integral · 0.7
YearPublicationVenuePosition
2026 Offline Behavioral Data Selection
abstract
Behavioral cloning is a widely adopted approach for offline policy learning from expert demonstrations. However, the large scale of offline behavioral datasets often results in computationally intensive training when used in downstream tasks. In this paper, we uncover the striking data saturation in offline behavioral data: policy performance rapidly saturates when trained on a small fraction of the dataset. We attribute this effect to the weak alignment between policy performance and test loss, revealing substantial room for improvement through data selection. To this end, we propose a simple yet effective method, Stepwise Dual Ranking (SDR), which extracts a compact yet informative subset from large-scale offline behavioral datasets. SDR is build on two key principles: (1) stepwise clip, which prioritizes early-stage data; and (2) dual ranking, which selects samples with both high action-value rank and low state-density rank. Extensive experiments and ablation studies on D4RL benchmarks demonstrate that SDR significantly enhances data selection for offline behavioral data. The code is available at https://github.com/LeavesLei/stepwise_dual_ranking.
Shiye Lei, Zhihao Cheng, Dacheng Tao
KDD (1)1
2025 Image Captions are Natural Prompts for Training Data Synthesis
Shiye Lei, Hao Chen 0011, Sen Zhang 0006, Bo Zhao 0015, Dacheng Tao
Int. J. Comput. Vis.1
2025 Attentive Learning Facilitates Generalization of Neural Networks
abstract
This article studies the generalization of neural networks (NNs) by examining how a network changes when trained on a training sample with or without out-of-distribution (OoD) examples. If the network's predictions are less influenced by fitting OoD examples, then the network learns attentively from the clean training set. A new notion, dataset-distraction stability, is proposed to measure the influence. Extensive CIFAR-10/100 experiments on the different VGG, ResNet, WideResNet, ViT architectures, and optimizers show a negative correlation between the dataset-distraction stability and generalizability. With the distraction stability, we decompose the learning process on the training set into multiple learning processes on the subsets of drawn from simpler distributions, i.e., distributions of smaller intrinsic dimensions (IDs), and furthermore, a tighter generalization bound is derived. Through attentive learning, miraculous generalization in deep learning can be explained and novel algorithms can also be designed.
Shiye Lei, Fengxiang He, Dacheng Tao
IEEE Trans. Neural Networks Learn. Syst.1
2025 Understanding Deep Learning via Decision Boundary
abstract
This article discovers that the neural network (NN) with lower decision boundary (DB) variability has better generalizability. Two new notions, algorithm DB variability and -data DB variability, are proposed to measure the DB variability from the algorithm and data perspectives. Extensive experiments show significant negative correlations between the DB variability and the generalizability. From the theoretical view, two lower bounds based on algorithm DB variability are proposed and do not explicitly depend on the sample size. We also prove an upper bound of order based on data DB variability. The bound is convenient to estimate without the requirement of labels and does not explicitly depend on the network size which is usually prohibitively large in deep learning.
Shiye Lei, Fengxiang He, Yancheng Yuan, Dacheng Tao
IEEE Trans. Neural Networks Learn. Syst.1
2024 Offline Behavior Distillation
abstract
Massive reinforcement learning (RL) data are typically collected to train policies offline without the need for interactions, but the large data volume can cause training inefficiencies. To tackle this issue, we formulate offline behavior distillation (OBD), which synthesizes limited expert behavioral data from sub-optimal RL data, enabling rapid policy learning. We propose two naive OBD objectives, DBC and PBC, which measure distillation performance via the decision difference between policies trained on distilled data and either offline data or a near-expert policy. Due to intractable bi-level optimization, the OBD objective is difficult to minimize to small values, which deteriorates PBC by its distillation performance guarantee with quadratic discount complexity $\mathcal{O}(1/(1-\gamma)^2)$. We theoretically establish the equivalence between the policy performance and action-value weighted decision difference, and introduce action-value weighted PBC (Av-PBC) as a more effective OBD objective. By optimizing the weighted decision difference, Av-PBC achieves a superior distillation guarantee with linear discount complexity $\mathcal{O}(1/(1-\gamma))$. Extensive experiments on multiple D4RL datasets reveal that Av-PBC offers significant improvements in OBD performance, fast distillation convergence speed, and robust cross-architecture/optimizer generalization.
Shiye Lei, Sen Zhang 0006, Dacheng Tao
NeurIPS1
2024 A Comprehensive Survey of Dataset Distillation
abstract
Deep learning technology has developed unprecedentedly in the last decade and has become the primary choice in many application domains. This progress is mainly attributed to a systematic collaboration in which rapidly growing computing resources encourage advanced algorithms to deal with massive data. However, it has gradually become challenging to handle the unlimited growth of data with limited computing power. To this end, diverse approaches are proposed to improve data processing efficiency. Dataset distillation, a dataset reduction method, addresses this problem by synthesizing a small typical dataset from substantial data and has attracted much attention from the deep learning community. Existing dataset distillation methods can be taxonomized into meta-learning and data matching frameworks according to whether they explicitly mimic the performance of target data. Although dataset distillation has shown surprising performance in compressing datasets, there are still several limitations such as distilling high-resolution data or data with complex label spaces. This paper provides a holistic understanding of dataset distillation from multiple aspects, including distillation frameworks and algorithms, factorized dataset distillation, performance comparison, and applications. Finally, we discuss challenges and promising directions to further promote future studies on dataset distillation.
Shiye Lei, Dacheng Tao
IEEE Trans. Pattern Anal. Mach. Intell.1
2023 Spectral complexity-scaled generalisation bound of complex-valued neural networks
abstract
Complex-valued neural networks (CVNNs) have been widely applied in various fields, primarily in signal processing and image recognition. Few studies have focused on the generalisation of CVNNs, although it is vital to ensure the performance of CVNNs on unseen data. This study is the first to prove a generalisation bound for complex-valued neural networks. The bounds increase as the spectral complexity increases, with the dominant factor being the product of the spectral norms of the weight matrices. Furthermore, this work provides a generalisation bound for CVNNs trained on sequential data, which is also affected by the spectral complexity. Theoretically, these bounds are derived using the Maurey Sparsification Lemma and Dudley entropy integral. We conducted empirical experiments on various datasets including MNIST, ashionMNIST, CIFAR-10, CIFAR-100, Tiny ImageNet, and IMDB by training complex-valued convolutional neural networks. The Spearman rank-order correlation coefficient and the corresponding p-values on these datasets provide strong proof of the statistically significant correlation between the spectral complexity of a network and its generalisation ability, as measured by the spectral norm product of the weight matrices. The code is available at https://github.com/LeavesLei/cvnn_generalization.
Fengxiang He, Shiye Lei, Dacheng Tao
Artif. Intell.3
2022 Accelerating temporal action proposal generation via high performance computing
Tian Wang 0002, Shiye Lei, Youyou Jiang, Chang Choi, Hichem Snoussi, Guangcun Shan
Frontiers Comput. Sci.2