EDBT 2026 Demo / reviewers in the wild / expert
Weihao Yang
dblp:243/2892
· DBLP profile ↗
6ranked-venue papers
0as first author
6since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 5 · 5 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | BirdMoE: Reducing Communication Costs for Mixture-of-Experts Training Using Load-Aware Bi-random QuantizationabstractMixture-of-Experts (MoE) model parallelism is prevalent in training Large Language Models (e.g., ChatGPT). However, the intensive all-to-all collective communication of the MoE layer’s intermediate computing results substantially degrades MoE training efficiency. In this paper, we propose BirdMoE, a novel load-aware communication compression technique with Bi-random quantization for MoE training with two core modules. Specifically, BirdMoE employs a lightweight Random Quantization (RQ) with expectation invariance property to efficiently map the floating-point intermediate computing results into integers while maintaining the MoE training quality. Additionally, BirdMoE utilizes a Mixed Precision (MP) strategy to dynamically balance the communication loads among expert nodes, significantly improving all-to-all communication efficiency for the MoE training system. Experiments on four typical MoE training tasks demonstrate that BirdMoE achieves higher $4.06 \times- 10.44 \times$ total communication compression ratios and $1.18 \times-5.27 \times$ training speedup compared with the state-of-the-art compression techniques while maintaining the MoE training quality. Donglei Wu, Weihao Yang, Xiangyu Zou, Jinda Jia, Dingwen Tao, Wen Xia, Zhihong Tian 0001 |
DAC | 2 |
| 2025 | A Novel Clustering Method for 2-D Low-Field NMR Spectra Working on Geological Fluid Parameter EstimationabstractLow-field nuclear magnetic resonance (LF-NMR) technology provides robust technical support for geophysical applications, including reservoir explorationand oil logging analysis. The inversion spectra of LF-NMR signals, particularly in two-dimensional (2D) form, reveal crucial geological information such as porosity, permeability, and other essential geologic parameters. However, the acquisition of geological parameters in geoscience relies on the scientific analysis, interpretation, and division of the inversion results, affecting the accuracy of the detection results. In LF-NMR, geophysical parameters are obtained by interpreting the fluid type and estimating the saturation of the inverted spectrum. Therefore, it is essential to develop corresponding qualitative and quantitative classification strategies, especially in cases involving fluid component aliasing. To address these geophysical issues comprehensively, we have created an improved fuzzy clustering algorithm using the local direction centrality (Fuzzy-CDC) based on 2D LF-NMR parameter spectra to observe the distribution characteristics of saturation for samples with overlapping states. Additionally, a fuzzy membership degree was introduced to enhance qualitative and quantitative abilities in assigning saturations to fluid components. To validate the clustering capability of this approach, we conducted simulation and actual experiments on four-phase fluids and water-gasoline two-phase fluids, respectively, comparing the results with conventional clustering methods. The improved method exhibits superior abilities and provided precise saturation estimation, yielding a relative error of 7.40% in simulation and 11.04% in actual experiments. In conclusion, our research significantly enhanced the analysis capabilities of 2D LF-NMR geophysical parameters while demonstrating the potential for pore fluid assessment and component classification in field geological exploration. Weihao Yang, Xiaoxue Lin, Tingting Lin 0001 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2024 | FedComp: A Federated Learning Compression Framework for Resource-Constrained Edge Computing DevicesabstractTop-K sparsification-based compression techniques are popular and powerful for reducing communication costs in federated learning (FL). However, existing Top-K sparsification-based compression methods suffer from two critical issues that severely hinder their implementation, particularly in the context of FL, which often involves a vast number of resource-constrained devices: 1) the low compressibility of the Top-K parameter’s indexes significantly limits the overall compression ratio (CR) and 2) the residual accumulation techniques used to maintain the model quality consume huge memory resources. To address these issues, we propose a novel FL compression framework, named FedComp, for deep neural networks (DNNs). FedComp achieves a higher communication CR while maintaining comparable model quality at low memory cost. Specifically, FedComp incorporates the following three key components: 1) a tensor-wise index-sharing mechanism that greatly reduces the index proportion by sharing one index among multiple elements of the tensor; 2) a fine-grained parameters packing strategy that reduces the transmission of duplicate value and index by considering their properties, thereby further reducing the overall communication cost; and 3) a residual compressor that significantly reduces memory cost by enhancing the compressibility of floating-point residuals and achieving a high CR with a lossless encoding scheme. Experiments on mainstream machine learning (ML) tasks with different DNN structures and datasets demonstrate that our proposed FedComp outperforms the state-of-the-art FL compression algorithms by achieving a higher communication CR of up to$28.5\times $while reducing memory costs by$21.04\times $–$50.59\times $on the local residual model, without degrading FL training performance. Donglei Wu, Weihao Yang, Haoyu Jin, Xiangyu Zou, Wen Xia, Binxing Fang |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2024 | BIRD+: Design of a Lightweight Communication Compressor for Resource-Constrained Distribution Learning PlatformsabstractThe Top-K sparsification-based compression framework is extensively explored for reducing communication costs in distributed learning. However, we identified several issues with existing Top-K sparsification-based compression methods: (i) The limited compressibility of the Top-K parameter's indexes critically restricts the overall communication compression ratio; (ii) Several time-consuming compression operations significantly offset the benefits of communication compression; (iii) The use of error feedback techniques to maintain model quality results in a high memory footprint consumption. To solve these issues, we propose BIRD, a lightweight tensor-wiseBi-Random samplingstrategy with an expectation invariance property. Specifically, BIRD applies a tensor-wiseindex sharingmechanism that reduces the index proportion by allowing multiple tensor elements to share a single index, thus improving the overall compression ratio. Additionally, BIRD replaces the time-consuming Top-K sorting with a fasterBi-Random samplingstrategy based on the aforementionedindex sharingmechanism, significantly reducing compression overheads; Moreover, BIRD establishes anexpectation invarianceproperty into theBi-Random samplingto ensure an approximate unbiased representation for the$L_1$-norm of the sampled tensors, effectively maintaining the model quality without incurring extra memory costs. We further optimize BIRD to BIRD+ by introducing the uniform distribution-based sampling and Gamma correction on the tensor-wise sampling process, achieving a more flexibly adjustment of the sparsity with better convergence performance. Experimental evaluations across multiple conventional distributed learning tasks demonstrate that compared to state-of-the-art approaches, BIRD+ achieves higher communication compression ratios up to 36.2$\times$and higher computation throughput up to 149.6$\times$while maintaining the model quality without incurring extra memory costs. Donglei Wu, Weihao Yang, Xiangyu Zou, Dingwen Tao, Wen Xia, Binxing Fang |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2023 | BIRD: A Lightweight and Adaptive Compressor for Communication-Efficient Distributed Learning Using Tensor-wise Bi-Random SamplingabstractTop-K sparsification-based compression framework is widely employed to reduce communication costs in distributed learning. However, we have identified several issues with existing Top-K sparsification-based compression methods that severely impede their deployment in resource-constrained devices: (i) the limited compressibility of the Top-K parameter’s indexes, which critically restricts the overall communication compression ratio; (ii) several time-consuming compression operations significantly negate the benefits of communication compression; (iii) the high memory footprint consumption associated with error feedback techniques used to maintain model quality.To address these issues, we propose a lightweight tensor-wise Bi-Random sampling strategy with expectation invariance property called BIRD, which achieves higher compression ratios at lower computational overheads while maintaining a comparable model quality without additional memory costs. Specifically, BIRD applies a tensor-wise index sharing mechanism that substantially reduces the proportion of the index by allowing multiple tensor elements to share a single index, thus improving the overall compression ratio. Additionally, BIRD replaces the time-consuming Top-K sorting with a faster Bi-Random sampling strategy based on the aforementioned index sharing mechanism, thereby reducing the computational costs of compression; Moreover, BIRD establishes an expectation invariance property into the above Bi-Random sampling to ensure an unbiased representation for the L1-norm of the sampled tensors, effectively maintaining the model quality without incurring extra memory costs.Experiments on multiple mainstream machine learning (ML) tasks demonstrate that compared to state-of-the-art methods, our proposed BIRD achieves 1.3×-31.1× higher compression ratio at lower time overheads with O(N) complexity while maintaining the model quality without incurring extra memory costs. Donglei Wu, Weihao Yang, Xiangyu Zou, Wen Xia |
ICCD | 2 |
| 2023 | Smart-DNN+: A Memory-efficient Neural Networks Compression Framework for the Model InferenceabstractDeep Neural Networks (DNNs) have achieved remarkable success in various real-world applications. However, running a Deep Neural Network (DNN) typically requires hundreds of megabytes of memory footprints, making it challenging to deploy on resource-constrained platforms such as mobile devices and IoT. Although mainstream DNNs compression techniques such as pruning, distillation, and quantization can reduce the memory overhead of model parameters during DNN inference, they suffer from three limitations: (i) low model compression ratio for the lightweight DNN structures with little redundancy, (ii) potential degradation in model inference accuracy, and (iii) inadequate memory compression ratio is attributable to ignoring the layering property of DNN inference. To address these issues, we propose a lightweight memory-efficient DNN inference framework called Smart-DNN+, which significantly reduces the memory costs of DNN inference without degrading the model quality. Specifically, ① Smart-DNN+ applies a layerwise binary-quantizer with a remapping mechanism to greatly reduce the model size by quantizing the typical floating-point DNN weights of 32-bit to the 1-bit signs layer by layer. To maintain model quality, ② Smart-DNN+ employs a bucket-encoder to keep the compressed quantization error by encoding the multiple similar floating-point residuals into the same integer bucket IDs. When running the compressed DNN in the user’s device, ③ Smart-DNN+ utilizes a partially decompressing strategy to greatly reduce the required memory overhead by first loading the compressed DNNs in memory and then dynamically decompressing the required materials for model inference layer by layer. Experimental results on popular DNNs and datasets demonstrate that Smart-DNN+ achieves lower 0.17%–0.92% memory costs at lower runtime overheads compared with the states of the art without degrading the inference accuracy. Moreover, Smart-DNN+ potentially reduces the inference runtime up to 2.04× that of conventional DNN inference workflow. Donglei Wu, Weihao Yang, Xiangyu Zou, Wen Xia, Zhenbo Hu, Weizhe Zhang, Binxing Fang |
ACM Trans. Archit. Code Optim. | 2 |