Donglei Wu

dblp:228/2096 · DBLP profile ↗
← Back
10ranked-venue papers
7as first author
10since 2021 · last 2026
0000-0003-0358-0533ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 8 · 5 first-author · 8 since 2021Artificial intelligence and machine learning · 2 · 2 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 first-author · 2 since 2021
YearPublicationVenuePosition
2026 RCMoE: A Communication-Efficient Random Compression Framework for Resource-Constrained Mixture-of-Experts Training
abstract
Mixture-of-Experts (MoE) architecture with experts parallelism scales LLMs efficiently by activating only a subset of experts per input, avoiding proportional training costs. However, the intensive and heterogeneous communication substantially hinders the efficiency and scalability of MoE training in the resource-constrained scenario. Existing communication compression techniques fall short in MoE training due to: (i) Intensive training amplifies compression overhead, compromising training efficiency; (ii) Accumulated compression errors propagate through the network, degrading training quality. In this paper, we propose RCMoE, a communication-efficient Random Compression framework for MoE training with two core modules: (1) Local-Stochastic Quantization compresses the all-to-all communication by stochastically quantizing each row of the expert's intermediate computing results in parallel, effectively improving the compression efficiency and reducing compression error; (2) Probabilistic Thresholding Sparsification compresses the all-reduce communication by probabilistically sampling large gradients at high probability, thereby reducing the computational complexity and maintaining the convergence efficiency. Experiments on four typical MoE training tasks prove that RCMoE achieves higher 5.9x-8.1x total communication compression ratios and 1.3x-10.1x training speedup compared with the state-of-the-art compression techniques while maintaining the MoE training accuracy.
Donglei Wu, Jinglei Tan, Jinda Jia, Guangming Tan, Dingwen Tao, Wen Xia, Zhihong Tian 0001
AAAI1
2025 BirdMoE: Reducing Communication Costs for Mixture-of-Experts Training Using Load-Aware Bi-random Quantization
abstract
Mixture-of-Experts (MoE) model parallelism is prevalent in training Large Language Models (e.g., ChatGPT). However, the intensive all-to-all collective communication of the MoE layer’s intermediate computing results substantially degrades MoE training efficiency. In this paper, we propose BirdMoE, a novel load-aware communication compression technique with Bi-random quantization for MoE training with two core modules. Specifically, BirdMoE employs a lightweight Random Quantization (RQ) with expectation invariance property to efficiently map the floating-point intermediate computing results into integers while maintaining the MoE training quality. Additionally, BirdMoE utilizes a Mixed Precision (MP) strategy to dynamically balance the communication loads among expert nodes, significantly improving all-to-all communication efficiency for the MoE training system. Experiments on four typical MoE training tasks demonstrate that BirdMoE achieves higher $4.06 \times- 10.44 \times$ total communication compression ratios and $1.18 \times-5.27 \times$ training speedup compared with the state-of-the-art compression techniques while maintaining the MoE training quality.
Donglei Wu, Weihao Yang, Xiangyu Zou, Jinda Jia, Dingwen Tao, Wen Xia, Zhihong Tian 0001
DAC1
2024 FedComp: A Federated Learning Compression Framework for Resource-Constrained Edge Computing Devices
abstract
Top-K sparsification-based compression techniques are popular and powerful for reducing communication costs in federated learning (FL). However, existing Top-K sparsification-based compression methods suffer from two critical issues that severely hinder their implementation, particularly in the context of FL, which often involves a vast number of resource-constrained devices: 1) the low compressibility of the Top-K parameter’s indexes significantly limits the overall compression ratio (CR) and 2) the residual accumulation techniques used to maintain the model quality consume huge memory resources. To address these issues, we propose a novel FL compression framework, named FedComp, for deep neural networks (DNNs). FedComp achieves a higher communication CR while maintaining comparable model quality at low memory cost. Specifically, FedComp incorporates the following three key components: 1) a tensor-wise index-sharing mechanism that greatly reduces the index proportion by sharing one index among multiple elements of the tensor; 2) a fine-grained parameters packing strategy that reduces the transmission of duplicate value and index by considering their properties, thereby further reducing the overall communication cost; and 3) a residual compressor that significantly reduces memory cost by enhancing the compressibility of floating-point residuals and achieving a high CR with a lossless encoding scheme. Experiments on mainstream machine learning (ML) tasks with different DNN structures and datasets demonstrate that our proposed FedComp outperforms the state-of-the-art FL compression algorithms by achieving a higher communication CR of up to$28.5\times $while reducing memory costs by$21.04\times $–$50.59\times $on the local residual model, without degrading FL training performance.
Donglei Wu, Weihao Yang, Haoyu Jin, Xiangyu Zou, Wen Xia, Binxing Fang
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.1
2024 BIRD+: Design of a Lightweight Communication Compressor for Resource-Constrained Distribution Learning Platforms
abstract
The Top-K sparsification-based compression framework is extensively explored for reducing communication costs in distributed learning. However, we identified several issues with existing Top-K sparsification-based compression methods: (i) The limited compressibility of the Top-K parameter's indexes critically restricts the overall communication compression ratio; (ii) Several time-consuming compression operations significantly offset the benefits of communication compression; (iii) The use of error feedback techniques to maintain model quality results in a high memory footprint consumption. To solve these issues, we propose BIRD, a lightweight tensor-wiseBi-Random samplingstrategy with an expectation invariance property. Specifically, BIRD applies a tensor-wiseindex sharingmechanism that reduces the index proportion by allowing multiple tensor elements to share a single index, thus improving the overall compression ratio. Additionally, BIRD replaces the time-consuming Top-K sorting with a fasterBi-Random samplingstrategy based on the aforementionedindex sharingmechanism, significantly reducing compression overheads; Moreover, BIRD establishes anexpectation invarianceproperty into theBi-Random samplingto ensure an approximate unbiased representation for the$L_1$-norm of the sampled tensors, effectively maintaining the model quality without incurring extra memory costs. We further optimize BIRD to BIRD+ by introducing the uniform distribution-based sampling and Gamma correction on the tensor-wise sampling process, achieving a more flexibly adjustment of the sparsity with better convergence performance. Experimental evaluations across multiple conventional distributed learning tasks demonstrate that compared to state-of-the-art approaches, BIRD+ achieves higher communication compression ratios up to 36.2$\times$and higher computation throughput up to 149.6$\times$while maintaining the model quality without incurring extra memory costs.
Donglei Wu, Weihao Yang, Xiangyu Zou, Dingwen Tao, Wen Xia, Binxing Fang
IEEE Trans. Parallel Distributed Syst.1
2023 BIRD: A Lightweight and Adaptive Compressor for Communication-Efficient Distributed Learning Using Tensor-wise Bi-Random Sampling
abstract
Top-K sparsification-based compression framework is widely employed to reduce communication costs in distributed learning. However, we have identified several issues with existing Top-K sparsification-based compression methods that severely impede their deployment in resource-constrained devices: (i) the limited compressibility of the Top-K parameter’s indexes, which critically restricts the overall communication compression ratio; (ii) several time-consuming compression operations significantly negate the benefits of communication compression; (iii) the high memory footprint consumption associated with error feedback techniques used to maintain model quality.To address these issues, we propose a lightweight tensor-wise Bi-Random sampling strategy with expectation invariance property called BIRD, which achieves higher compression ratios at lower computational overheads while maintaining a comparable model quality without additional memory costs. Specifically, BIRD applies a tensor-wise index sharing mechanism that substantially reduces the proportion of the index by allowing multiple tensor elements to share a single index, thus improving the overall compression ratio. Additionally, BIRD replaces the time-consuming Top-K sorting with a faster Bi-Random sampling strategy based on the aforementioned index sharing mechanism, thereby reducing the computational costs of compression; Moreover, BIRD establishes an expectation invariance property into the above Bi-Random sampling to ensure an unbiased representation for the L1-norm of the sampled tensors, effectively maintaining the model quality without incurring extra memory costs.Experiments on multiple mainstream machine learning (ML) tasks demonstrate that compared to state-of-the-art methods, our proposed BIRD achieves 1.3×-31.1× higher compression ratio at lower time overheads with O(N) complexity while maintaining the model quality without incurring extra memory costs.
Donglei Wu, Weihao Yang, Xiangyu Zou, Wen Xia
ICCD1
2023 Smart-DNN+: A Memory-efficient Neural Networks Compression Framework for the Model Inference
abstract
Deep Neural Networks (DNNs) have achieved remarkable success in various real-world applications. However, running a Deep Neural Network (DNN) typically requires hundreds of megabytes of memory footprints, making it challenging to deploy on resource-constrained platforms such as mobile devices and IoT. Although mainstream DNNs compression techniques such as pruning, distillation, and quantization can reduce the memory overhead of model parameters during DNN inference, they suffer from three limitations: (i) low model compression ratio for the lightweight DNN structures with little redundancy, (ii) potential degradation in model inference accuracy, and (iii) inadequate memory compression ratio is attributable to ignoring the layering property of DNN inference. To address these issues, we propose a lightweight memory-efficient DNN inference framework called Smart-DNN+, which significantly reduces the memory costs of DNN inference without degrading the model quality. Specifically, ① Smart-DNN+ applies a layerwise binary-quantizer with a remapping mechanism to greatly reduce the model size by quantizing the typical floating-point DNN weights of 32-bit to the 1-bit signs layer by layer. To maintain model quality, ② Smart-DNN+ employs a bucket-encoder to keep the compressed quantization error by encoding the multiple similar floating-point residuals into the same integer bucket IDs. When running the compressed DNN in the user’s device, ③ Smart-DNN+ utilizes a partially decompressing strategy to greatly reduce the required memory overhead by first loading the compressed DNNs in memory and then dynamically decompressing the required materials for model inference layer by layer. Experimental results on popular DNNs and datasets demonstrate that Smart-DNN+ achieves lower 0.17%–0.92% memory costs at lower runtime overheads compared with the states of the art without degrading the inference accuracy. Moreover, Smart-DNN+ potentially reduces the inference runtime up to 2.04× that of conventional DNN inference workflow.
Donglei Wu, Weihao Yang, Xiangyu Zou, Wen Xia, Zhenbo Hu, Weizhe Zhang, Binxing Fang
ACM Trans. Archit. Code Optim.1
2023 Design of a Quantization-Based DNN Delta Compression Framework for Model Snapshots and Federated Learning
abstract
Deep neural networks (DNNs) have achieved remarkable success in many fields. However, large-scale DNNs also bring storage costs when storing snapshots for preventing clusters’ frequent failures or incur significant communication overheads when transmitting DNNs in the Federated Learning (FL). Recently, several approaches, such as Delta-DNN and LC-Checkpoint, aim to reduce the size of DNNs’ snapshot storage by compressing the difference between two neighboring versions of the DNNs (a.k.a., delta). However, we observe that existing approaches, applying traditional global lossy quantization techniques in DNN's delta compression, can not fully exploit the data similarity since the parameters’ value ranges vary among layers. To fully explore the similarity of the delta model and improve the compression ratio, we propose a quantization-based local-sensitive delta compression approach, named QD-Compressor, by developing a layer-based local-sensitive quantization scheme and error feedback mechanism. Specifically, the quantizers and number of quantization bits are adaptive among layers based on the value distribution and weighted entropy of the delta's parameters. To avoid quantization error degrading the performance of the restored model, an alternative error feedback mechanism is designed to dynamically correct the quantization error during the training process. Experiments on multiple popular DNNs and datasets show that QD-Compressor obtains a higher 7×-40× compression ratio in the model snapshot compression scenario than the state-of-the-art approaches. Additionally, QD-Compressor achieves an 11×-15× compression ratio to the residual model of the Federated Learning compression scenario.
Haoyu Jin, Donglei Wu, Xiangyu Zou, Sian Jin, Dingwen Tao, Qing Liao 0001, Wen Xia
IEEE Trans. Parallel Distributed Syst.2
2022 SmartIdx: Reducing Communication Cost in Federated Learning by Exploiting the CNNs Structures
abstract
Top-k sparsification method is popular and powerful forreducing the communication cost in Federated Learning(FL). However, according to our experimental observation, it spends most of the total communication cost on the index of the selected parameters (i.e., their position informa-tion), which is inefficient for FL training. To solve this problem, we propose a FL compression algorithm for convolution neural networks (CNNs), called SmartIdx, by extending the traditional Top-k largest variation selection strategy intothe convolution-kernel-based selection, to reduce the proportion of the index in the overall communication cost and thusachieve a high compression ratio. The basic idea of SmartIdx is to improve the 1:1 proportion relationship betweenthe value and index of the parameters to n:1, by regarding the convolution kernel as the basic selecting unit in parameter selection, which can potentially deliver more informationto the parameter server under the limited network traffic. Tothis end, a set of rules are designed for judging which kernel should be selected and the corresponding packaging strategies are also proposed for further improving the compressionratio. Experiments on mainstream CNNs and datasets show that our proposed SmartIdx performs 2.5×−69.2× higher compression ratio than the state-of-the-art FL compression algorithms without degrading model performance.
Donglei Wu, Xiangyu Zou, Haoyu Jin, Wen Xia, Binxing Fang
AAAI1
2021 Smart-DNN: Efficiently Reducing the Memory Requirements of Running Deep Neural Networks on Resource-constrained Platforms
abstract
Deep neural networks (DNNs) have gained considerable attention in various real-world applications due to their strong performance in representation learning. However, running a DNN needs tremendous memory resources, which significantly restricts DNN from being applicable on resource-constrained platforms (e.g., IoT, mobile devices, etc.). Lightweight DNNs can accommodate the characteristics of mobile devices, but the hardware resources of mobile or IoT devices are extremely limited, and the resource consumption of lightweight models needs to be further reduced. However, the current neural network compression approaches (i.e., pruning, quantization, knowledge distillation, etc.) works poorly on the lightweight DNNs, which are already simplified. In this paper, we present a novel framework called Smart-DNN, which can efficiently reduce the memory requirements of running DNNs on resource-constrained platforms. Specifically, we slice a neural network into several segments and use SZ error-bounded lossy compression to compress each segment separately while keeping the network structure unchanged. When running a network, we first store the compressed network into memory and then partially decompress the corresponding part layer by layer. According to experimental results on four popular lightweight DNNs (usually used in resource-constrained platforms), Smart-DNN achieves memory saving of 1/10∼1/5, while slightly sacrificing inference accuracy and unchanging the neural network structure with accepted extra runtime overhead.
Zhenbo Hu, Xiangyu Zou, Wen Xia, Weizhe Zhang, Donglei Wu
ICCD6
2021 QD-Compressor: a Quantization-based Delta Compression Framework for Deep Neural Networks
abstract
Deep neural networks (DNNs) have achieved remarkable success in many fields. Large-scale DNNs also bring storage challenges when storing snapshots for preventing clusters’ frequent failures, and bring massive internet traffic when dispatching or updating DNNs for resource-constrained devices (e.g., IoT devices, mobile phones). Several approaches are aiming to compress DNNs. The Recent work, Delta-DNN, notices high similarity existed in DNNs and thus calculates differences between them for improving the compression ratio.However, we observe that Delta-DNN, applying traditional global lossy quantization technique in calculating differences of two neighboring versions of the DNNs, can not fully exploit the data similarity between them for delta compression. This is because the parameters’ value ranges (and also the delta data in Delta-DNN) are varying among layers in DNNs, which inspires us to propose a local-sensitive quantization scheme: the quantizers are adaptive to parameters’ local value ranges in layers. Moreover, instead of quantizing differences of DNNs in Delta-DNN, our approach quantizes DNNs before calculating differences to make the differences more compressible. Besides, we also propose an error feedback mechanism to reduce DNNs’ accuracy loss caused by the lossy quantization.Therefore, we design a novel quantization-based delta compressor called QD-Compressor, which calculates the lossy differences between epochs of DNNs for saving storage cost of backing up DNNs’ snapshots and internet traffic of dispatching DNNs for resource-constrained devices. Experiments on several popular DNNs and datasets show that QD-Compressor obtains a compression ratio of 2.4× ~ 31.5× higher than the state-of-the-art approaches while well maintaining the model’s test accuracy.
Donglei Wu, Haoyu Jin, Xiangyu Zou, Wen Xia, Xiaojia Huang
ICCD2