Yongmei Lei

dblp:08/4194 · DBLP profile ↗
← Back
15ranked-venue papers
0as first author
8since 2021 · last 2025
0009-0008-8010-5545ORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 7 · 3 since 2021Artificial intelligence and machine learning · 4 · 3 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2025 The Fast Inertial ADMM optimization framework for distributed machine learning
Yongmei Lei
Future Gener. Comput. Syst.4
2024 Distributed and Correlation Group-based Sure Independence Screening and Sparsifying Operator
abstract
SISSO (Sure Independence Screening and Sparsifying Operator) is a data-driven method that combines symbolic regression and compressed sensing to effectively identify the most representative features in large-scale datasets. However, when constructing high-dimensional models, SISSO retains too many redundant features when establishing high-dimensional models, leading to increased computational costs and accuracy issues, especially in small-sample datasets. A correlation group screening strategy (CGS) is designed to remove feature redundancy in SISSO, reducing the feature space for sparse operations and further lowering computational costs. Based on this strategy, a CG-SISSO method is implemented. Additionally, to address the accuracy issues caused by differences in multi-subsystem data within small datasets, the distributed parallel CG-SISSO model is proposed, which involves balancing the feature selection across multiple subsystem subsets. Experiments were conducted on three datasets with six targets each, demonstrating that CG-SISSO can significantly reduce the algorithm’s time complexity by one-tenth while ensuring stable accuracy. When building models for a multi-system dataset of perovskite, the DBCG-SISSO model successfully improved the fitting accuracy.
Yongmei Lei
IJCNN2
2024 A Fast ADMM Framework for Training Deep Neural Networks Without Gradients
abstract
Stochastic gradient descent (SGD) and its many variants are widely used algorithms for training deep neural networks (DNN). However, SGD has some unavoidable drawbacks, including vanishing gradients, great sensitivity to inputs, and lack of theoretical guarantees. To address these drawbacks, the Alternating Direction Method of Multipliers (ADMM) has been proposed as an effective alternative to gradient-based methods with gradient-free properties. It has been successfully used to train deep neural networks. Nevertheless, existing ADMM-based approaches face several challenges, including slow convergence, complex update processes, and being single-task oriented. In this paper, we propose a novel optimization framework for deep learning via fast ADMM (fdlADMM) to address these drawbacks. The fdlADMM algorithm combines the benefits of the fast ADMM method with the characteristics of the DNN structure to enhance model convergence speed and improve performance. This algorithm introduces novel ideas and methods to the domain of deep learning. Additionally, the algorithm introduces a novel update process and utilizes a quadratic approximation to efficiently solve the sub-problem, thereby simplifying the update process and reducing computation time. We conduct extensive experiments on two distinct tasks, namely image classification and node classification, utilizing four benchmark datasets. The goal is to validate the effectiveness and efficiency of the proposed fdlADMM algorithm across varied domains and tasks. The results show that the fdlADMM algorithm converges significantly faster, achieves better accuracy, and maintains a very competitive training speed compared to other state-of-the-art ADMM-based optimizers.
Xueting Wen, Yongmei Lei
IJCNN2
2024 S-ADMM: Optimizing Machine Learning on Imbalanced Datasets Based on ADMM
abstract
Nowadays, addressing large-scale machine learning problems using high-performance computing (HPC) clusters has gained significant importance. The alternating direction method of multipliers (ADMM) is widely used in machine learning for solving optimization problems on clusters. However, ADMM's performance is greatly affected by imbalanced datasets on the HPC cluster. In this paper, we propose the distributed shunt ADMM with a new adaptive penalty method based on the hybrid MPI/OpenMP programming model (S-ADMM). The proposed shunt strategy chooses different sub-problem optimization algorithms to improve the accuracy with the imbalanced datasets. Additionally, we design a novel adaptive penalty parameter method and improve a sub-problem optimization algorithm for S-ADMM. The adaptive penalty parameter method accelerates the algorithm's convergence and the sub-problem optimization algorithm improves the training efficiency of ADMM. Moreover, S-ADMM reduces communication cost by exchanging parameters among nodes using the MPI and saves calculation time by parallel computation within nodes via OpenMP threads. For the SVM classification problem, experiments conducted on the Tianhe-2 supercomputing platform show that S-ADMM has competitive running efficiency and up to 43% accuracy improvement compared to existing distributed ADMM implemented with pure MPI or MPI/OpenMP on imbalanced datasets.
Yixiang Shen, Yongmei Lei, Qinnan Qiu
SMC2
2023 PSRA-HGADMM: A Communication Efficient Distributed ADMM Algorithm
abstract
Among distributed machine learning algorithms, the global consensus alternating direction method of multipliers (ADMM) has attracted much attention because it can effectively solve large-scale optimization problems. However, the high communication cost slows its convergence and limits scalability. To solve the problem, we propose a hierarchical grouping ADMM algorithm (PSRA-HGADMM) with a novel Ring-Allreduce communication model in this paper. Firstly, we optimize the parameter exchange of the ADMM algorithm and implement the global consensus ADMM algorithm in the decentralized architecture. Secondly, to improve the communication efficiency of the distributed system, we propose a novel Ring-Allreduce communication model (PSR-Allreduce) based on the idea of parameter server architecture. Finally, a Worker-Leader-Group generator (WLG) framework is designed to solve the problem of inconsistency of cluster nodes. This framework combines hierarchical parameter aggregation and adopts the grouping strategy to improve the scalability of the distributed system. Experiments show that PSRA-HGADMM has better convergence performance and better scalability than ADMMLib and AD-ADMM. Compared with ADMMLib, the overall communication cost of PSRA-HGADMM is reduced by 32%.
Yongwen Qiu, Yongmei Lei
ICPP2
2023 A Communication Efficient ADMM-based Distributed Algorithm Using Two-Dimensional Torus Grouping AllReduce
abstract
Abstract Large-scale distributed training mainly consists of sub-model parallel training and parameter synchronization. With the expansion of training workers, the efficiency of parameter synchronization will be affected. To tackle this problem, we first propose 2D-TGA, a grouping AllReduce method based on the two-dimensional torus topology. This method synchronizes the model parameters by grouping and makes full use of bandwidth. Secondly, we propose a distributed algorithm, 2D-TGA-ADMM, which combines the 2D-TGA with the alternating direction method of multipliers (ADMM). It focuses on sub-model training and reduces the wait time among workers in the synchronization process. Finally, experimental results on the Tianhe-2 supercomputing platform show that compared with the $${\mathtt {MPI\_Allreduce}}$$ MPI _ Allreduce , the 2D-TGA could shorten the synchronization wait time by $$33\%$$ 33 % .
Yongmei Lei, Cunlu Peng
Data Sci. Eng.2
2023 Communication-efficient ADMM-based distributed algorithms for sparse training
Yongmei Lei, Yongwen Qiu, Lingfei Lou
Neurocomputing2
2021 HSAC-ALADMM: an asynchronous lazy ADMM algorithm based on hierarchical sparse allreduce communication
Dongxia Wang 0003, Yongmei Lei, Jinyang Xie
J. Supercomput.2
2020 A Dynamic Scheduling Strategy of ADMM Sub-problem Optimization Algorithm Based on Hierarchical Structure
Jiawei Ji, Yongmei Lei, Shenghong Jiang
ICA3PP (1)2
2019 Collaborative Computing of Urban Built-Up Area Identification from Remote Sensing Image
Yongmei Lei, Jun-Juan Zhao
CollaborateCom3
2019 ADMMLIB: A Library of Communication-Efficient AD-ADMM for Distributed Machine Learning
Jinyang Xie, Yongmei Lei
NPC2
2018 Fast Communication Structure for Asynchronous Distributed ADMM Under Unbalance Process Arrival Pattern
Shuqing Wang, Yongmei Lei
ICANN (1)2
2013 BSP-based support vector regression machine parallel framework
abstract
In this paper, we propose a BSP-Based Support Vector Regression Machine Parallel Framework which can implement the most of distributed Support Vector Regression Machine algorithms. The major difference in these algorithms is the network topology among distributed nodes. Therefore, we adopt the Bulk Synchronous Parallel model to solve the strongly connected graph problem in exchanging support vectors among distributed nodes. Besides, we introduce the dynamic algorithms that it can change the strongly connected graph among SVR distributed nodes in every BSP's super-step. The performance of this framework has been analyzed and evaluated with KDD99 data and four DPSVR algorithms with different topology on the high-performance computer. The results proved that the framework can implement the most of distributed SVR algorithms and keep the performance of original algorithm.
Yongmei Lei
ICIS2
2009 Generation of Web Knowledge Flow for Personalized Services
abstract
Personalized services provide specific services to satisfy userpsilas real-time requirements. How to effectively discover and organize proper Web resources is a key issue. Based on Web knowledge flow which is used to semantically organize and represent Web resources that are recommended to user, this paper presents a 3-way strategy of finding proper resources for a user in multi-user environment. With this strategy, generation of Web Knowledge flow can be accomplished based on collaborative user and semantics link network. First, according to the similarity degree between active user and collaborative user, collaborative user is refined to strict kind and loose kind. Secondly, userpsilas browsing sequence of Web resources is presented to find collaborative users and implemented in a 3-D coordinate space. Thirdly, average information entropy of semantic relationship between Web resources is presented to evaluate the similarity of userspsila browsing feature. The experimental results demonstrate the validity of the method. It can be seen that the proposed method has a brilliant perspective in the applications of Web personalized services.
Jie Yu 0009, Xiangfeng Luo, Feiyue Ye, Yongmei Lei
ISPA4
2005 Security solution to manufacturing grid usage scenarios
abstract
Manufacturing grid consists of a collection of heterogeneous resources across multiple administrative domains with the intent of providing enterprises to realize resource sharing and collaboration. A comprehensive set of manufacturing grid usage scenarios is presented and analyzed with regard to security requirements such as authentication, authorization, integrity, nonrepudition and confidentiality. The main value of these scenarios is to increase the awareness of security issues in manufacturing grid which are different from computational grid. We take the unique security requirements of manufacturing grid into consideration and design a corresponding security solution. The implementation within the project of SMVPN (Shanghai Manufacturing Virtual Private Network) proves that the solution is workable and valuable.
Hongxia Cai, Minglun Fang, Yongmei Lei
CCGRID4