Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Xupeng Shi

dblp:251/8550 · DBLP profile ↗
← Back
4ranked-venue papers
1as first author
3since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 4 · 1 first-author · 3 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
3 papers
Efficient and distributed learning · 73% Optimization for machine learning · 14% Learning theory · 11%

Topics — the 8 heaviest of 9, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Efficient and distributed learning
model compression
1.322025
Computation and Memory-Efficient Model Compression with Gradient Reweighting · NeurIPS 2025
Revisiting Parameter Sharing for Automatic Neural Channel Number Search · NeurIPS 2020
Machine learning › Efficient and distributed learning
distributed training
0.912025
Computation and Memory-Efficient Model Compression with Gradient Reweighting · NeurIPS 2025
Machine learning › Optimization for machine learning › gradient-based optimization
gradient reweighting
0.912025
Computation and Memory-Efficient Model Compression with Gradient Reweighting · NeurIPS 2025
Machine learning › Efficient and distributed learning › model compression
pruning
0.912025
Computation and Memory-Efficient Model Compression with Gradient Reweighting · NeurIPS 2025
Machine learning › Learning theory
margin-based learning
0.712023
DynaMS: Dyanmic Margin Selection for Efficient Deep Learning · ICLR 2023
Machine learning › Efficient and distributed learning › automated machine learning
neural architecture search
0.412020
Revisiting Parameter Sharing for Automatic Neural Channel Number Search · NeurIPS 2020
Machine learning › Efficient and distributed learning
parameter sharing
0.412020
Revisiting Parameter Sharing for Automatic Neural Channel Number Search · NeurIPS 2020
Machine learning › Deep learning architectures and training
convolutional neural network
0.112020
Revisiting Parameter Sharing for Automatic Neural Channel Number Search · NeurIPS 2020

Methods — techniques the papers use, named apart from their topics

preconditioning · 0.9gradient estimation · 0.9clipping · 0.9dynamic margin selection · 0.7affine parameter sharing · 0.4
YearPublicationVenuePosition
2025 Computation and Memory-Efficient Model Compression with Gradient Reweighting
abstract
Pruning is a commonly employed technique for deep neural networks (DNNs) aiming at compressing the model size to reduce computational and memory costs during inference. In contrast to conventional neural networks, large language models (LLMs) pose a unique challenge regarding pruning efficiency due to their substantial computational and memory demands. Existing methods, particularly optimization-based ones, often require considerable computational resources in gradient estimation because they cannot effectively leverage weight sparsity of the intermediate pruned network to lower compuation and memory costs in each iteration. The fundamental challenge lies in the need to frequently instantiate intermediate pruned sub-models to achieve these savings, a task that becomes infeasible even for moderately sized neural networks. To this end, this paper proposes a novel pruning method for DNNs that is both computationally and memory-efficient. Our key idea is to develop an effective reweighting mechanism that enables us to estimate the gradient of the pruned network in current iteration via reweigting the gradient estimated on an outdated intermediate sub-model instantiated at an earlier stage, thereby significantly reducing model instantiation frequency. We further develop a series of techniques, e.g., clipping and preconditioning matrix, to reduce the variance of gradient estimation and stabilize the optimization process. We conducted extensive experimental validation across various domains. Our approach achieves 50\% sparsity and a 1.58$\times$ speedup in forward pass on Llama2-7B model with only 6 GB of memory usage, outperforming state-of-the-art methods with respect to both perplexity and zero-shot performance. As a by-product, our method is highly suited for distributed sparse training and can achieve a 2 $\times$ speedup over the dense distributed baselines.
Yuesen Liao, Binrui Wu, Yuquan Zhou, Xupeng Shi, Dongsheng Jiang
NeurIPS5
2023 DynaMS: Dyanmic Margin Selection for Efficient Deep Learning
Jingwei Zhuo, Xupeng Shi, Lixing Gong, Tong Tao, Pengzhang Liu, Yongjun Bao, Weipeng Yan
ICLR4
2022 Finding Dynamics Preserving Adversarial Winning Tickets
abstract
Modern deep neural networks (DNNs) are vulnerable to adversarial attacks and adversarial training has been shown to be a promising method for improving the adversarial robustness of DNNs. Pruning methods have been considered in adversarial context to reduce model capacity and improve adversarial robustness simultaneously in training. Existing adversarial pruning methods generally mimic the classical pruning methods for natural training, which follow the ’training, pruning, fine-tuning’ three stages pipeline. We observe that such pruning methods do not necessarily preserve the dynamics of dense networks, making it potentially hard to be fine-tuned to compensate the accuracy degradation in pruning. Based on recent works of neural tangent kernel (NTK), we systematically study the dynamics of adversarial training and prove the existence of trainable sparse sub-network at initialization which can be trained to be adversarial robust from scratch. This theoretically verifies the lottery ticket hypothesis in adversarial context and we refer such sub-network structure as adversarial winning ticket (AWT). We also show empirical evidences that AWT preserves the dynamics of adversarial training and achieve equal performance as dense adversarial training.
Xupeng Shi, A. Adam Ding, Yuan Gao 0015
AISTATS1
2020 Revisiting Parameter Sharing for Automatic Neural Channel Number Search
abstract
Recent advances in neural architecture search inspire many channel number search algorithms~(CNS) for convolutional neural networks. To improve searching efficiency, parameter sharing is widely applied, which reuses parameters among different channel configurations. Nevertheless, it is unclear how parameter sharing affects the searching process. In this paper, we aim at providing a better understanding and exploitation of parameter sharing for CNS. Specifically, we propose affine parameter sharing~(APS) as a general formulation to unify and quantitatively analyze existing channel search algorithms. It is found that with parameter sharing, weight updates of one architecture can simultaneously benefit other candidates. However, it also results in less confidence in choosing good architectures. We thus propose a new strategy of parameter sharing towards a better balance between training efficiency and architecture discrimination. Extensive analysis and experiments demonstrate the superiority of the proposed strategy in channel configuration against many state-of-the-art counterparts on benchmark datasets.
Haoli Bai, Jiaxiang Wu 0001, Xupeng Shi, Junzhou Huang, Irwin King, Michael R. Lyu, Jian Cheng 0001
NeurIPS4