EDBT 2026 Demo / reviewers in the wild / expert
Xupeng Shi
dblp:251/8550
· DBLP profile ↗
4ranked-venue papers
1as first author
3since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 4 · 1 first-author · 3 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
3 papers |
Efficient and distributed learning · 73% Optimization for machine learning · 14% Learning theory · 11% |
Topics — the 8 heaviest of 9, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Efficient and distributed learning
model compression |
1.3 | 2 | 2025 | Computation and Memory-Efficient Model Compression with Gradient Reweighting · NeurIPS 2025 Revisiting Parameter Sharing for Automatic Neural Channel Number Search · NeurIPS 2020 |
Machine learning › Efficient and distributed learning
distributed training |
0.9 | 1 | 2025 | Computation and Memory-Efficient Model Compression with Gradient Reweighting · NeurIPS 2025 |
Machine learning › Optimization for machine learning › gradient-based optimization
gradient reweighting |
0.9 | 1 | 2025 | Computation and Memory-Efficient Model Compression with Gradient Reweighting · NeurIPS 2025 |
Machine learning › Efficient and distributed learning › model compression
pruning |
0.9 | 1 | 2025 | Computation and Memory-Efficient Model Compression with Gradient Reweighting · NeurIPS 2025 |
Machine learning › Learning theory
margin-based learning |
0.7 | 1 | 2023 | DynaMS: Dyanmic Margin Selection for Efficient Deep Learning · ICLR 2023 |
Machine learning › Efficient and distributed learning › automated machine learning
neural architecture search |
0.4 | 1 | 2020 | Revisiting Parameter Sharing for Automatic Neural Channel Number Search · NeurIPS 2020 |
Machine learning › Efficient and distributed learning
parameter sharing |
0.4 | 1 | 2020 | Revisiting Parameter Sharing for Automatic Neural Channel Number Search · NeurIPS 2020 |
Machine learning › Deep learning architectures and training
convolutional neural network |
0.1 | 1 | 2020 | Revisiting Parameter Sharing for Automatic Neural Channel Number Search · NeurIPS 2020 |
Methods — techniques the papers use, named apart from their topics
preconditioning · 0.9gradient estimation · 0.9clipping · 0.9dynamic margin selection · 0.7affine parameter sharing · 0.4
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Computation and Memory-Efficient Model Compression with Gradient ReweightingabstractPruning is a commonly employed technique for deep neural networks (DNNs) aiming at compressing the model size to reduce computational and memory costs during inference. In contrast to conventional neural networks, large language models (LLMs) pose a unique challenge regarding pruning efficiency due to their substantial computational and memory demands. Existing methods, particularly optimization-based ones, often require considerable computational resources in gradient estimation because they cannot effectively leverage weight sparsity of the intermediate pruned network to lower compuation and memory costs in each iteration. The fundamental challenge lies in the need to frequently instantiate intermediate pruned sub-models to achieve these savings, a task that becomes infeasible even for moderately sized neural networks. To this end, this paper proposes a novel pruning method for DNNs that is both computationally and memory-efficient. Our key idea is to develop an effective reweighting mechanism that enables us to estimate the gradient of the pruned network in current iteration via reweigting the gradient estimated on an outdated intermediate sub-model instantiated at an earlier stage, thereby significantly reducing model instantiation frequency. We further develop a series of techniques, e.g., clipping and preconditioning matrix, to reduce the variance of gradient estimation and stabilize the optimization process. We conducted extensive experimental validation across various domains. Our approach achieves 50\% sparsity and a 1.58$\times$ speedup in forward pass on Llama2-7B model with only 6 GB of memory usage, outperforming state-of-the-art methods with respect to both perplexity and zero-shot performance. As a by-product, our method is highly suited for distributed sparse training and can achieve a 2 $\times$ speedup over the dense distributed baselines. Yuesen Liao, Binrui Wu, Yuquan Zhou, Xupeng Shi, Dongsheng Jiang |
NeurIPS | 5 |
| 2023 | DynaMS: Dyanmic Margin Selection for Efficient Deep Learning
Jingwei Zhuo, Xupeng Shi, Lixing Gong, Tong Tao, Pengzhang Liu, Yongjun Bao, Weipeng Yan |
ICLR | 4 |
| 2022 | Finding Dynamics Preserving Adversarial Winning TicketsabstractModern deep neural networks (DNNs) are vulnerable to adversarial attacks and adversarial training has been shown to be a promising method for improving the adversarial robustness of DNNs. Pruning methods have been considered in adversarial context to reduce model capacity and improve adversarial robustness simultaneously in training. Existing adversarial pruning methods generally mimic the classical pruning methods for natural training, which follow the ’training, pruning, fine-tuning’ three stages pipeline. We observe that such pruning methods do not necessarily preserve the dynamics of dense networks, making it potentially hard to be fine-tuned to compensate the accuracy degradation in pruning. Based on recent works of neural tangent kernel (NTK), we systematically study the dynamics of adversarial training and prove the existence of trainable sparse sub-network at initialization which can be trained to be adversarial robust from scratch. This theoretically verifies the lottery ticket hypothesis in adversarial context and we refer such sub-network structure as adversarial winning ticket (AWT). We also show empirical evidences that AWT preserves the dynamics of adversarial training and achieve equal performance as dense adversarial training. Xupeng Shi, A. Adam Ding, Yuan Gao 0015 |
AISTATS | 1 |
| 2020 | Revisiting Parameter Sharing for Automatic Neural Channel Number SearchabstractRecent advances in neural architecture search inspire many channel number search algorithms~(CNS) for convolutional neural networks. To improve searching efficiency, parameter sharing is widely applied, which reuses parameters among different channel configurations. Nevertheless, it is unclear how parameter sharing affects the searching process. In this paper, we aim at providing a better understanding and exploitation of parameter sharing for CNS. Specifically, we propose affine parameter sharing~(APS) as a general formulation to unify and quantitatively analyze existing channel search algorithms. It is found that with parameter sharing, weight updates of one architecture can simultaneously benefit other candidates. However, it also results in less confidence in choosing good architectures. We thus propose a new strategy of parameter sharing towards a better balance between training efficiency and architecture discrimination. Extensive analysis and experiments demonstrate the superiority of the proposed strategy in channel configuration against many state-of-the-art counterparts on benchmark datasets. Haoli Bai, Jiaxiang Wu 0001, Xupeng Shi, Junzhou Huang, Irwin King, Michael R. Lyu, Jian Cheng 0001 |
NeurIPS | 4 |