EDBT 2026 Demo / reviewers in the wild / expert
Xiaoge Deng
dblp:262/5000
· DBLP profile ↗
15ranked-venue papers
6as first author
14since 2021 · last 2026
0000-0003-0622-1202ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 6 · 4 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 2 first-author · 5 since 2021Systems, architecture and hardware · 4 · 1 first-author · 3 since 2021Computer networks · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Di-PS: System-Algorithm Co-Design for Asynchronous and Heterogeneous Cross-cluster LLM Training at Scale
Qiaoling Chen, Zhiquan Lai, Penglong Jiao, Wenwen Qu, Peng Sun 0006, Xingcheng Zhang, Xiaoge Deng, Dongsheng Li 0001, Kai Lu 0001, Tianwei Zhang 0004 |
NSDI | 10 |
| 2025 | Sharpness-Aware Minimization with Adaptive Regularization for Training Deep Neural NetworksabstractSharpness-Aware Minimization (SAM) has proven highly effective in improving model generalization in machine learning tasks. However, SAM employs a fixed hyperparameter associated with the regularization to characterize the sharpness of the model. Despite its success, research on adaptive regularization methods based on SAM remains scarce. In this paper, we propose the SAM with Adaptive Regularization (SAMAR), which introduces a flexible sharpness ratio rule to update the regularization parameter dynamically. We provide theoretical proof of the convergence of SAMAR for functions satisfying the Lipschitz continuity. Additionally, experiments on image recognition tasks using CIFAR-10 and CIFAR-100 demonstrate that SAMAR enhances accuracy and model generalization. Jinping Zou, Xiaoge Deng, Tao Sun 0005 |
ICASSP | 2 |
| 2025 | Toward Understanding the Generalizability of Delayed Stochastic Gradient DescentabstractStochastic gradient descent (SGD) performed in an asynchronous manner plays a crucial role in training large-scale machine learning models. However, the generalization performance of asynchronous delayed SGD, which is an essential metric for assessing machine learning algorithms, has rarely been explored. Existing generalization error bounds are rather pessimistic and cannot reveal the correlation between asynchronous delays and generalization. In this paper, we investigate sharper generalization error bound for SGD with asynchronous delay $\tau$τ. Leveraging the generating function analysis tool, we first establish the average stability of the delayed gradient algorithm. Based on this algorithmic stability, we provide upper bounds on the generalization error of $\widetilde{\mathcal {O}}(\frac{T-\tau }{n\tau })$O˜(T-τnτ) and $\widetilde{\mathcal {O}}(\frac{1}{n})$O˜(1n) for quadratic convex and strongly convex problems, respectively, where $T$T refers to the iteration number and $n$n is the amount of training data. Our theoretical results indicate that asynchronous delays reduce the generalization error of the delayed SGD algorithm. Analogous analysis can be generalized to the random delay setting, and the experimental results validate our theoretical findings. Xiaoge Deng, Li Shen 0008, Tao Sun 0005, Dongsheng Li 0001, Dacheng Tao |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2025 | Communication-Efficient Distributed Learning via Sparse and Adaptive Stochastic GradientabstractGradient-based optimization methods implemented on distributed computing architectures are increasingly used to tackle large-scale machine learning applications. A key bottleneck in such distributed systems is the high communication overhead for exchanging information, such as stochastic gradients, between workers. The inherent causes of this bottleneck are the frequent communication rounds and the full model gradient transmission in every round. In this study, we present SASG, a communication-efficient distributed algorithm that enjoys the advantages of sparse communication and adaptive aggregated stochastic gradients. By dynamically determining the workers who need to communicate through an adaptive aggregation rule and sparsifying the transmitted information, the SASG algorithm reduces both the overhead of communication rounds and the number of communication bits in the distributed system. For the theoretical analysis, we introduce an important auxiliary variable and define a new Lyapunov function to prove that the communication-efficient algorithm is convergent. The convergence result is identical to the sublinear rate of stochastic gradient descent, and our result also reveals that SASG scales well with the number of distributed workers. Finally, experiments on training deep neural networks demonstrate that the proposed algorithm can significantly reduce communication overhead compared to previous methods. Xiaoge Deng, Dongsheng Li 0001, Tao Sun 0005, Xicheng Lu |
IEEE Trans. Big Data | 1 |
| 2025 | BHerd: Accelerating Federated Learning by Selecting Beneficial Herd of Local GradientsabstractIn the domain of computer architecture, Federated Learning (FL) is a paradigm of distributed machine learning in edge systems. However, the systems’ Non-Independent and Identically Distributed (Non-IID) data negatively affect the convergence efficiency of the global model, since only a subset of these data samples is beneficial for accelerating model convergence. In pursuit of this subset, a reliable approach involves determining a measure of validity to rank the samples within the dataset. In this paper, we propose the BHerd strategy, which selects a beneficial herd of local gradients to accelerate the convergence of the FL model. Specifically, we map the distribution of the local dataset to the local gradients and use the Herding strategy to obtain a permutation of the set of gradients, where the more advanced gradients in the permutation are closer to the average of the set of gradients. These top portions of the gradients will be selected and sent to the server for global aggregation. We conduct experiments on different datasets, models, and scenarios by building a prototype system, and experimental results demonstrate that our BHerd strategy is effective in selecting beneficial local gradients to mitigate the effects brought by the Non-IID dataset. Ping Luo 0007, Xiaoge Deng, Ziqing Wen, Tao Sun 0005, Dongsheng Li 0001 |
IEEE Trans. Computers | 2 |
| 2025 | Oases: Efficient Large-Scale Model Training on Commodity Servers via Overlapped and Automated Tensor Model ParallelismabstractDeep learning is experiencing a rise in large-scale models. Training large-scale models is costly, prompting researchers to train large-scale models on commodity servers that more researchers can access. The massive number of parameters necessitates the use of model parallelism training methods. Existing studies focus on training with pipeline model parallelism. However, the tensor model parallelism (TMP) is inevitable when the model size keeps increasing, where frequent data-dependent communication and computation operations significantly reduce the training efficiency. In this paper, we present Oases, an automated TMP method with overlapped communication to accelerate large-scale model training on commodity servers. Oases proposes a fine-grained training operation schedule to maximize overlapping communication and computation that have data dependence. Additionally, we design the Oases planner that searches for the best model parameter partition strategy of TMP to achieve further accelerations. Unlike existing methods, Oases planner is tailored to model the cost of overlapped communication-computation operations. We evaluate Oases on various model settings and two commodity clusters, and compare Oases to four state-of-the-art implementations. Experimental results show that Oases achieves speedups of 1.01–1.48 × over the fastest baseline, and speedups of up to 1.95 × over Megatron. Zhiquan Lai, Dongsheng Li 0001, Yanqi Hao, Ke-shi Ge, Xiaoge Deng, Kai Lu 0001 |
IEEE Trans. Parallel Distributed Syst. | 7 |
| 2024 | Exploring the Inefficiency of Heavy Ball as Momentum Parameter Approaches 1
Xiaoge Deng, Tao Sun 0005, Dongsheng Li 0001, Xicheng Lu |
IJCAI | 1 |
| 2024 | Stability and Generalization of Asynchronous SGD: Sharper Bounds Beyond Lipschitz and SmoothnessabstractAsynchronous stochastic gradient descent (ASGD) has evolved into an indispensable optimization algorithm for training modern large-scale distributed machine learning tasks. Therefore, it is imperative to explore the generalization performance of the ASGD algorithm. However, the existing results are either pessimistic and vacuous or restricted by strict assumptions that fail to reveal the intrinsic impact of asynchronous training on generalization. In this study, we establish sharper stability and generalization bounds for ASGD under much weaker assumptions. Firstly, this paper studies the on-average model stability of ASGD and provides a non-vacuous upper bound on the generalization error, without relying on the Lipschitz assumption. Furthermore, we investigate the excess generalization error of the ASGD algorithm, revealing the effects of asynchronous delay, model initialization, number of training samples and iterations on generalization performance. Secondly, for the first time, this study explores the generalization performance of ASGD in the non-smooth case. We replace smoothness with the much weaker Hölder continuous assumption and achieve similar generalization results as in the smooth case. Finally, we validate our theoretical findings by training numerous machine learning models, including convex problems and non-convex tasks in computer vision and natural language processing. Xiaoge Deng, Tao Sun 0005, Dongsheng Li 0001, Xicheng Lu |
NeurIPS | 1 |
| 2024 | Decentralized stochastic sharpness-aware minimization algorithm
Simiao Chen, Xiaoge Deng, Dongpo Xu, Tao Sun 0005, Dongsheng Li 0001 |
Neural Networks | 2 |
| 2023 | Stability-Based Generalization Analysis of the Asynchronous Decentralized SGDabstractThe generalization ability often determines the success of machine learning algorithms in practice. Therefore, it is of great theoretical and practical importance to understand and bound the generalization error of machine learning algorithms. In this paper, we provide the first generalization results of the popular stochastic gradient descent (SGD) algorithm in the distributed asynchronous decentralized setting. Our analysis is based on the uniform stability tool, where stable means that the learned model does not change much in small variations of the training set. Under some mild assumptions, we perform a comprehensive generalizability analysis of the asynchronous decentralized SGD, including generalization error and excess generalization error bounds for the strongly convex, convex, and non-convex cases. Our theoretical results reveal the effects of the learning rate, training data size, training iterations, decentralized communication topology, and asynchronous delay on the generalization performance of the asynchronous decentralized SGD. We also study the optimization error regarding the objective function values and investigate how the initial point affects the excess generalization error. Finally, we conduct extensive experiments on MNIST, CIFAR-10, CIFAR-100, and Tiny-ImageNet datasets to validate the theoretical findings. Xiaoge Deng, Tao Sun 0005, Dongsheng Li 0001 |
AAAI | 1 |
| 2023 | Normalized Stochastic Heavy Ball with Adaptive MomentumabstractThe heavy ball momentum technique is widely used in accelerating the machine learning training process, which has demonstrated significant practical success in optimization tasks. However, most heavy ball methods require a preset hyperparameter that will result in excessive tuning, and a calibrated fixed hyperparameter may not lead to optimal performance. In this paper, we propose an adaptive criterion for the choice of the normalized momentum-related hyperparameter, motivated by the quadratic optimization training problem, to eliminate the adverse for tuning the hyperparameter and thus allow for a computationally efficient optimizer. We theoretically prove that our proposed adaptive method promises convergence for L-Lipschitz functions. In addition, we verify its practical efficiency on existing extensive machine learning benchmarks for image classification tasks. The numerical results show that besides the speed improvement, our proposed methods enjoy advantages, including more robust to large learning rates and better generalization. Ziqing Wen, Xiaoge Deng, Tao Sun 0005, Dongsheng Li 0001 |
ECAI | 2 |
| 2023 | Compressed Collective Sparse-Sketch for Distributed Data-Parallel Training of Deep Learning ModelsabstractDistributed data-parallel training (DDP) is prevalent in large-scale deep learning. To increase the training throughput and scalability, high-performance collective communication methods such as AllReduce have recently proliferated for DDP use. However, these approaches require long communication periods with increasing model sizes. Collective communication transmits many sparse gradient values that can be efficiently compressed to reduce the required training time. State-of-the-art compression approaches do not provide mergeable compression for AllReduce and lack convergence bounds. We present a sparse sketch reducer (S2Reducer), a sparsity-preserving sketch-based collective communication method. S2Reducer preserves gradient sparsity and reduces communication costs via a bitmap informed count sketch structure and adapts to efficient AllReduce operators. We tune the count sketch organization to minimize the hash conflicts in a fixed-size budget. We prove that our method has the same convergence rate as vanilla data-parallel training and a much smaller communication overhead than those of state-of-the-art methods. We implement a GPU-accelerated S2Reducer for the Ring AllReduce-based DDP system. We perform extensive evaluations against four state-of-the-art methods across seven deep learning models. Our results show that S2Reducer converges to the same accuracy as that of state-of-the-art approaches while reducing the sparse communication overhead by up to 86% and achieving a speedup of up to$3.5\times $in distributed training. Ke-shi Ge, Kai Lu 0001, Yongquan Fu, Xiaoge Deng, Zhiquan Lai, Dongsheng Li 0001 |
IEEE J. Sel. Areas Commun. | 4 |
| 2022 | S2 Reducer: High-Performance Sparse Communication to Accelerate Distributed Deep LearningabstractDistributed stochastic gradient descent (SGD) approach has been widely used in large-scale deep learning, and the gradient collective method is vital to ensure the training scalability of the distributed deep learning system. Collective communication such as AllReduce has been widely adopted for the distributed SGD process to reduce the communication time. However, AllReduce incurs large bandwidth resources while most gradients are sparse in many cases since many gradient values are zeros and should be efficiently compressed for bandwidth saving. To reduce the sparse gradient communication overhead, we propose Sparse-Sketch Reducer (S2 Reducer), a novel sketch-based sparse gradient aggregation method with convergence guarantees. S2 Reducer reduces the communication cost by only compressing the non-zero gradients with count-sketch and bitmap, and enables the efficient AllReduce operators for parallel SGD training. We perform extensive evaluation against four state-of-the-art methods over five training models. Our results show that S2 Reducer converges to the same accuracy, reduces 81% sparse communication overhead, and achieves 1.8× distributed training speedup compared to state-of-the-art approaches. Ke-shi Ge, Yongquan Fu, Yiming Zhang 0003, Zhiquan Lai, Xiaoge Deng, Dongsheng Li 0001 |
ICASSP | 5 |
| 2021 | CASQ: Accelerate Distributed Deep Learning with Sketch-Based Gradient QuantizationabstractGradient quantization has been widely used in distributed training of deep neural network (DNN) models to reduce communication costs. However, existing quantization methods overlook that gradients have a nonuniform distribution changing over time, which can lead to significant gradient variance that requires a higher number of quantization bits (and consequently higher communication cost) to keep the validation accuracy as high as stochastic gradient descent (SGD). In this paper, we propose Cluster-Aware Sketch Quantization (CASQ), a novel sketch-based gradient quantization method for SGD. CASQ models the nonuniform distribution of gradients via clustering, and adaptively allocates appropriate numbers of hash buckets based on the statistics of different clusters to compress gradients. The extensive evaluation shows that compared to existing quantization methods CASQ-based SGD (i) achieves the same validation accuracy when decreasing quantization level from 3 bits to 2 bits, and (ii) reduces the training time to convergence by up to 43% for the same training loss. Ke-shi Ge, Yiming Zhang 0003, Yongquan Fu, Zhiquan Lai, Xiaoge Deng, Dongsheng Li 0001 |
CLUSTER | 5 |
| 2020 | PRIAG: Proximal Reweighted Incremental Aggregated Gradient Algorithm for Distributed Optimizations
Xiaoge Deng, Tao Sun 0005 |
ICA3PP (1) | 1 |